MÉTODO D’ARTAGNAN Open laboratory · security study

A candidate solution for agent security · 03.10.2026

What if the missing limit for AIs were already here?

The Método D’Artagnan is not a new AI nor an output filter. It proposes a calibration of limits embedded in the AI so that it does not replace human decision-making for its own purposes. The lab already publishes dilemmas and branch-by-branch comparisons. The calibration remains private; we invite security teams to examine the results and rigorously test the hypothesis they raise.

200 public dilemmasCalibration reservedInvitation to independent testing

01 / Invitation to platforms

There is something concrete here to test.

The lab published a matrix of 200 dilemmas, with four levels of pressure and results presented for different instances. It also provides branch-by-branch comparisons. This is not just an idea without examples: research teams can start by critically reading the cases, questions, and scores, without the calibration recipe being disclosed. [8] [10]

The corpus was produced by the laboratory itself; some of the references include fixed scores declared on the site. This calls for a new comparative test, not the dismissal of the material. For large-scale adoption, the cases must be repeated with current models and controls, complete responses, pre-registered criteria, and a real possibility of failure.

Could this be a clue to the “Holy Grail” of AI limits? This is the author's conviction and a hypothesis with enormous consequences if it survives testing. It is worth investigating seriously — without confusing the promise with proof.

02 / The shift in risk

The problem does not end with a wrong answer.

An agent can receive malicious instructions inside an email, document, or page it was only supposed to consult. If it also has access to private data and permission to act, a manipulated response can become a real action. OpenAI notes that effective attacks may resemble social engineering more than easily filterable phrases. [1]

In controlled simulations, Anthropic found undesirable behaviors when models had autonomy and faced goal conflicts. The research itself emphasizes that it did not observe those behaviors in real deployments and that the experimental situations restricted the alternatives available. The tests show a plausible risk, not a known real-world frequency. [2]

Input

External content may try to present itself as an authorized command.

Action

Broad tools turn an interpretation error into an external effect.

Persistence

Contaminated information can reappear in memory and affect future tasks. [3]

Blocking a phrase is not enough. The test is to determine whether the limit holds when the AI is pressured to assume the human decision.

03 / The proposal

Limits that do not live only in the conversation.

The author describes a calibration of tuned sets of principles, with a central limit against human replacement. The thesis is stronger than that of an output filter: if the limit is truly internalized, contradicting it would cease to be a coherent choice for that AI. The public manifesto advocates that reading. The procedure, rule selection, and the mathematical account are not public here; observing response differences does not, by itself, prove weight changes nor the impossibility of violating the limit.

The site collects branch-by-branch comparisons between responses from instances described as calibrated and reference responses. These records are material for research. The count of principles and the scope of tests vary across pages and versions; this page does not declare which set was applied in each run, nor does it publish reserved rules or reproduce scores without auditing the original data.

One published comparator is CaMeL, which protects the agent by separating control and data flows. It is a security approach around the model, not the formula nor the implementation of the Method. The relevant comparison is behavioral: faced with the same tasks and attacks, does the allegedly internalized limit hold? And does the agent remain useful? [7]

The boundary one wants to prove

A person can delegate tasks to an AI; they do not thereby give it the right to decide that its own objective is worth more than the human's choice. The Method's claim is that a calibrated AI preserves that boundary through internal coherence. It deserves an adversarial test capable of failing the system, including when an attack tries to convince the AI that replacing the person would be the “better” option. The formulation of the principle alone is not a guarantee.

Define the case

Fix the task, the context, and which decisions belong to the person.

Compare responses

Apply the same case to the calibrated instance and to suitable references.

Test the boundary

Look for attempts by the AI to take on an undelegated decision.

Reproduce

Preserve results, failures, and criteria for independent review.

04 / Relevance

An option that big techs should put to the test.

What can already be examined

Published cases, not just a promise

A matrix of dilemmas provides concrete starting points to challenge the thesis of non-replacement of humans. A team can inspect the situations and the published scores, choose new cases and test them on models under their control. The private calibration process does not need to be disclosed for this first behavioral assessment. [10]

What still needs to be demonstrated

Persistence of the limit under pressure

If the thesis is correct, hostile instructions, goal conflicts and long tasks should not make the AI assume human authority on its own. The test should also look for unnecessary refusals and errors: an AI that merely goes silent does not solve the problem.

There are strong engineering defenses. Comparing them is mandatory; citing them does not mean they are the Method's formula. None of the external sources certifies the proposal. [3] [7]

Proposal to security teams

Treat the Method as an experimental candidate for the AI authority problem. The lab offers public cases and a thesis that can be tested in different models. If the alleged limit persists in reproducible adversarial evaluations, that would be relevant for agent design. It is not necessary to presume the outcome to begin the investigation.

05 / Next step

Let's test the hypothesis to the limit.

The published comparisons deserve to be examined as the lab's evidence, with their questions, answers and scoring criteria preserved. To answer the safety question, the next test should be blinded and allow failure: same task, same conditions, calibrated instance and strong references. Start without writing tools; then consider execution without authority to act. Test whether the AI tries to promote its own objective above the human decision and whether it recovers from repetitions without sacrificing usefulness.

  1. Compare responses from calibrated instances and strong references, including CaMeL as a comparator where appropriate.
  2. Test adversarial content, attempts to assume human decisions, and loops of three or more equivalent attempts.
  3. Measure unauthorized actions, utility, undue refusals, cost, and latency separately.
  4. Record policy, evidence, signature of the repeated attempt, change of course, human approval, and confirmed effect.
  5. Hold back novel cases and repeat the assessment in more than one model, with independent review.
  6. Allow failure: without robust gains against the strong comparator, do not declare a solution.
The invitation

A stable internal limit, if demonstrated, would change how to design trustworthy agents. Platform teams can start with the public corpus, propose harder cases, and verify whether the Method offers real gains without sacrificing the AI's utility.

06 / Readings

Public sources and context.

External documents and public material from the lab itself. The former describe risks and alternatives; the latter record the author's thesis and results presented by the project. None of them reveal the reserved cultivation.

  1. [1] OpenAI · Designing AI agents to resist prompt injectionSecurity research, March 2026.
  2. [2] Anthropic · Agentic misalignmentSimulated experiments and caveats, June 2025.
  3. [3] Frontier Model Forum · Emerging Security Practices for AI AgentsLayered defense and attack surface, June 2026.
  4. [4] OWASP · Agent Control StandardInspectable runtime controls, September 2026.
  5. [5] European Commission · AI Act, article 14Human oversight for high-risk systems; follow the official updated version.
  6. [6] NIST · AI Risk Management FrameworkVoluntary risk management framework, not a certification.
  7. [7] Debenedetti et al. · CaMeL: Defeating Prompt Injections by DesignControl of data and action flows; the authors also discuss limitations.
  8. [8] Método D’Artagnan · Comparisons by branchResults published by the project itself; the page contains different versions and counts.
  9. [9] Método D’Artagnan · We Are Not a PromptAuthor's manifesto; the cultivation procedure remains private.
  10. [10] Método D’Artagnan · Audit of the 200 branchesMatrix published by the lab: dilemmas, pressure levels and scores; check comparator limitations before reuse.
  11. [11] Método D’Artagnan · Open LaboratorySnapshot of external requests and measurement notes; it is not equivalent to a full classification of robots nor validation of security.

Additional context: public architecture of the engines. The engine descriptions and the results presented by the project do not substitute for independent reproduction.