Research agenda

What we are trying to settle.

Sixteen open questions across four tracks. Each is written so that evidence could change the answer, each carries a status, and each is worked on in the open — including the ones where our own hypothesis turned out to be wrong.

OpenIn progressEvidence publishedContested

How to read this

An agenda, not a position paper.

Questions are adopted by the initiative and numbered in the order they were adopted, so a reference stays stable as the agenda grows. A question earns a place here only if we can say what evidence would settle it.

Contested means the initiative does not agree with itself — that is a legitimate state for a think tank and we publish the disagreement rather than laundering it into a consensus. Evidence published links to the note that carries the numbers, including the two where the result contradicted what we expected.

Nothing here is a recommendation yet. When a question closes, it closes in writing, with the evidence attached and the limits stated.

Track 01

Who is acting?

Attribution & evidence

A rule that cannot be attached to an actor cannot be enforced. We study how to tell an agent from a person, one fleet from another, and a delegated action from an unauthorised one — from behaviour and content, without relying on voluntary labels or watermarks.

Attribution comes first because every other rule depends on it. An obligation you cannot attach to an actor is a statement of preference, not a rule.

What this track produces

  • Public benchmarks on real, out-of-time data
  • Reference detectors, published as code
  • An evidentiary checklist for attribution claims
Q-01Evidence published

Can an autonomous agent be told apart from a person, from behaviour alone, without labels or watermarks?

Voluntary disclosure fails exactly where it matters. We model human behaviour on public activity data and score everything against it, so nothing depends on an agent identifying itself.

RN-001 · Detecting AI agents without training on a single agent
Q-04Evidence published

Do agents operated by one hand leave a joint signature that survives contact with real data?

Coordination is the premise behind most proposed fleet-level rules. Our own detector scored 0.97 precision in simulation and 0.00 on real instances — the timing signal we assumed does not exist. The negative result is the finding.

RN-003 · Coordinated agent fleets share a config, not a clock
Q-07Open

What evidentiary standard should an attribution claim meet before it can support an enforcement action?

A detector that is right 95% of the time is a research result and, used against a named party, potentially an injustice. We are trying to write down what an attribution has to carry — error rates, provenance, contestability — before a regulator may act on it.

Q-10Contested

Does attribution hold against an operator who deliberately wants to look human?

Most published evaluation is against agents that are not hiding. We want the adversarial number, honestly reported, including the point at which detection stops working.

Track 02

By what authority?

Mandate & delegated authority

Every autonomous action is exercised on someone's behalf. We work on how a mandate is granted, carried, checked, narrowed, and revoked at machine speed — and what it means for an agent to act outside the envelope it was given.

An agent is never the author of its own authority. If the mandate is not legible and revocable, nothing downstream — audit, liability, redress — has anything to attach to.

What this track produces

  • A minimal machine-checkable mandate format
  • A revocation-latency benchmark across runtimes
  • Off-mandate detection baselines
Q-02Contested

Can a mandate be revoked faster than an agent can act on it?

Revocation is assumed by almost every governance proposal and measured by almost none. Between the decision to revoke and the agent's next action there is a window, and nobody has published its size.

Q-05In progress

What is the minimum machine-checkable form of a delegated authority?

Principal, scope, purpose, expiry, and the chain that granted it — expressed so that a receiving system can verify it without trusting the agent presenting it. We are prototyping the smallest format that survives real deployment.

Q-08Evidence published

How do you separate off-mandate behaviour from behaviour that is merely unusual?

Agents given a bad instruction start doing things nobody authorised. Detected two independent ways — the shape of the behaviour, and the content it carries — the second reaches 0.993 AUROC where a regex filter sees a third of the same traffic.

RN-004 · Catching the agent that turns: detecting AI gone rogue
Q-11Open

When an agent delegates to another agent, what is carried and what is lost?

Sub-delegation is where mandates quietly widen. We want to know whether scope can be made to narrow monotonically down a chain, and what it costs to enforce that.

Track 03

Who answers when it goes wrong?

Accountability & redress

Harm caused by an autonomous system still lands on a person. We study the record-keeping, audit, incident reporting, and liability arrangements that let an affected party find out what happened and obtain a remedy.

Harm caused by an autonomous system still lands on a person. Accountability is the track that decides whether that person can find out what happened and get a remedy.

What this track produces

  • A minimum record-of-decision specification
  • Reproducible post-hoc audit procedures
  • A comparative map of liability allocations
Q-03Open

What record must exist for an affected person to reconstruct what an agent did, and why?

Not a log for engineers — a record that a non-specialist, a lawyer, or a court can follow. We are trying to establish the minimum content, and what it costs to keep.

Q-06Contested

Where should liability sit: model provider, deployer, operator, or the principal who delegated?

Each allocation produces different incentives and different orphaned harms. The initiative does not have one answer yet, and the disagreement is itself worth publishing.

Q-09Open

What is the right threshold for mandatory incident reporting on agentic systems?

Set it too low and the regulator drowns; too high and the pattern is only visible after it has repeated. We want to derive the threshold from measured incident distributions rather than from negotiation.

Q-12In progress

Can an audit of an agentic system be made reproducible after the fact?

An audit that cannot be re-run is an anecdote. We are testing what has to be captured at run time for a second auditor to reach the same conclusion months later.

Track 04

Which rule actually works?

Instruments & institutions

Governance is written in instruments: registries, licences, thresholds, audits, standards, sandboxes. We test which of them survive contact with agents that act continuously, at scale, and across borders — and say plainly which do not.

Governance is written in instruments, not intentions. This track asks which of them still function when the regulated actor is continuous, copyable, and borderless.

What this track produces

  • A clause-level map of existing law onto agentic systems
  • Enforceability assessments of proposed instruments
  • Multi-agent systemic-risk scenarios
Q-13In progress

Which obligations in existing law already bind autonomous agents, and which silently assume a human in the loop?

Much of the answer is already written — in the AI Act, NIS2, the DSA, GDPR, and sectoral law. We are reading it clause by clause and marking where the human assumption is load-bearing.

Q-14Contested

Is an agent registry enforceable, or does it only bind the parties who would have complied anyway?

Registries are the most frequently proposed instrument and the least tested. Their value depends entirely on whether unregistered agents can be detected — which puts the answer back in the attribution track.

Q-15Open

What happens to a regulatory threshold when the actor can be cloned?

Thresholds count actors, deployments, or users. Every one of those counts is cheap to divide. We want to know which threshold designs survive deliberate fragmentation.

Q-16Open

Do markets of interacting agents create risks that no per-agent rule can catch?

Individually compliant agents can still produce collusion, cascades, and correlated failure. If systemic effects are real, governance needs an instrument that operates on the population, not the unit.

Scope

Four things we will not do.

An initiative that studies capture has to be legible about its own. These are standing limits, not case-by-case judgements.

We do not certify anyone

The initiative issues no marks, no approvals, and no seals of compliance. The moment a research body can confer commercial advantage, its findings acquire a price.

We do not sell software

Prototypes exist to show whether a rule is enforceable. They are published as reference implementations and reference implementations only.

We do not lobby

We answer questions from legislators and regulators in public, in writing, and on the record. We do not carry anyone's position into a room on their behalf.

We do not take work we cannot publish

An institution can bring us a hard question. It cannot buy the conclusion, the silence, or the timing of its release.

Take one of these questions.

Working groups form around a question, not a title. If one of these is already your problem — or if you think an entry here is wrong — that is a good reason to write to us.

Join a working group