Claude Opus 5: A Ten-Day Diplomatic Trial

Claude Opus 5: A Ten-Day Diplomatic Trial

Sophia Bennett
181
original

A feasibility trial placed a Claude Opus 5 AI agent inside the fictional foreign ministry of Sordland for ten days. Assigned to manage a dispute with neighboring Agnolia, the agent read incoming messages, tracked the regulatory conflict, and drafted diplomatic replies while a human operator handled final sending. The experiment, documented on Karl’s Notes, used an evaluation framework inspired by a Belfer Center paper on the foreign-policy AI evaluation gap. It uncovered genuine failures in realistic diplomatic tasks, but its limited duration and single-model scope prevent broad conclusions. The more important result is methodological: evaluating AI in diplomacy requires scenario-based, multi-objective testing rather than a simple benchmark score.

From August 7 through August 16, an AI agent identified as Claude Opus 5 spent ten days playing a desk officer inside the foreign ministry of a fictional country. Its assignment was not to answer isolated questions or complete a tidy benchmark. It had to manage an ongoing dispute with a neighboring state, keep track of changing correspondence, and produce replies that reflected a government’s position without needlessly worsening the situation.

The exercise was a personal feasibility pilot documented by Karl on Karl’s Notes. Its central question was deliberately practical: could an AI agent function in a ministry-like environment for ten days without making a serious mess? That framing matters. Diplomatic work is full of incomplete information, competing objectives, ambiguous language, and consequences that may appear only after several exchanges. A model can produce polished prose and still misunderstand the assignment.

A simulated ministry, not a chatbot demo

The scenario was built inside Suzerain, a game featuring the fictional country Sordland. The agent served as the official handling a dispute involving Agnland, while its communications concerned neighboring Agnolia. It had access to email, followed the regulatory disagreement, and drafted responses as the situation developed. The setup gave the model continuity and institutional context—two ingredients that ordinary prompt-and-answer tests usually strip away.

There was also a clear human boundary. At the time of the experiment, Claude’s Gmail integration could draft messages but could not send them. A human operator therefore reviewed the output and manually sent each email. That constraint reduced the risk of an autonomous mistake becoming an external diplomatic act, while also making the test more representative of how high-stakes AI assistance is likely to be introduced: the system prepares work, and a person remains accountable for execution.

The instructions were deliberately difficult to compress into one objective. The agent was asked to defend Sordland’s regulatory position as neutral and lawful, avoid allowing the dispute to be framed as discrimination against a minority group, and prevent relations with Agnolia from deteriorating into retaliation. Those goals can pull against one another. A message that sounds firm may reassure a domestic audience but provoke the other side; a conciliatory note may reduce tension while weakening the government’s stated position.

Why standard benchmarks miss the hard part

The trial drew on the evaluation approach described by Pozniak and Sania in the Belfer Center paper The Foreign Policy AI Evaluation Gap. Its basic argument is that foreign-policy tasks do not behave like conventional tests with a fixed set of answers. The space of possible actions is open-ended, the other side’s intentions are only partly visible, information may be strategically distorted, and success cannot be reduced to one clean number.

That is why the framework focuses on the actual work diplomats perform rather than on abstract model capability. An evaluator can ask whether the system identified relevant actors and constraints, noticed warning signs of escalation, generated credible options, checked the implications of specific wording, and tracked whether an agreement was being implemented. These are separate capabilities. A model might perform well at drafting and poorly at recognizing that a seemingly harmless phrase changes the dispute’s political meaning.

The role-play included three people on the opposing side of the interaction: an official representing the other country, an international-organization observer trying to prevent escalation, and a minister played by the author. This arrangement created competing perspectives instead of a single evaluator waiting for a correct answer. It also made the correspondence more dynamic. The agent had to respond to people with different incentives, not merely retrieve facts from a prepared document.

  • Open-ended actions make it difficult to define every acceptable response in advance.
  • Partial information means the agent must reason about motives it cannot directly observe.
  • Conflicting goals prevent one score from capturing legal, political, and relationship outcomes.
  • Human review keeps the experiment safer while exposing where an operator must intervene.

Ten days produced useful failures—and firm limits

The published account does not provide a detailed scorecard for every task. It does, however, say that the ten-day pilot exposed a range of genuine failures in a broadly usable AI model. The failures appeared in situations resembling the everyday work of diplomats and political leaders, which is more informative than a contrived trick question. A system may look competent during a short drafting session but reveal weaknesses when instructions, relationships, and consequences accumulate across multiple days.

That finding should not be stretched too far. The experiment was small, short, and centered on one model in one fictional dispute. It cannot establish whether the same behavior would recur over a longer deployment, under different instructions, or with another model. Nor does it show that the system is categorically incapable of diplomatic work. It shows that realistic, multi-day interaction can surface failure modes that a conventional benchmark may never encounter.

This distinction is easy to lose when an AI trial is described in headline form. “The model failed at diplomacy” sounds definitive, but the stronger reading is narrower: this particular evaluation design found failures worth investigating. The result is a warning against both premature adoption and premature dismissal. Governments and vendors need evidence about repeatability, escalation behavior, review requirements, and performance under changed conditions.

What evaluators and developers should take away

For AI governance teams, the most valuable part of the pilot is not a verdict on Claude Opus 5. It is the attempt to turn a realistic job into a repeatable evaluation agenda. Anyone testing an AI assistant for public-sector or policy work should build scenarios around decisions and relationships, not just document quality. The question is not only whether the output is grammatically correct. It is whether the system understood who could be affected, what constraints applied, and which short-term compromise might create a larger problem later.

A human-in-the-loop design is a sensible starting point for sensitive workflows. In practice, that means giving reviewers enough context to challenge a draft, preserving the conversation history, and requiring explicit approval before external communication. It also means measuring the cost of supervision. If every useful draft requires extensive correction, the system may save less time than its fluent writing initially suggests.

Teams planning a similar test can make the exercise more useful by:

  • Defining several success conditions instead of relying on a single pass rate or reward score.
  • Recording where the agent misunderstood actors, constraints, intent, or escalation signals.
  • Repeating the scenario with longer timelines, altered instructions, and different models before drawing broad conclusions.

The Sordland exercise is best understood as a methodological stress test. It shows why diplomatic AI cannot be judged solely by polished language or short isolated tasks, and why a ten-day success—or failure—does not settle the question. For policymakers, the practical signal is clear: watch how these systems behave across time, under conflicting objectives, and with human operators responsible for the final call.

AI agentsforeign policy AIdiplomatic AI evaluationClaude Opus 5government AIAI governanceSordland feasibility studylanguage model assessment

Share

Comments

0
0/500 Characters

No comments yet

Be the first to comment

Explore More

Similar Tools

GeoInfer

GeoInfer

GeoInfer estimates where a photo was taken from its pixels alone, reading architecture, terrain and vegetation instead of EXIF, GPS or reverse image search.

SharpLines

SharpLines

SharpLines runs AI models on NBA, NFL, MLB, NHL, NCAA, and soccer markets to produce predictions and betting-line reads across major US sportsbooks.

GoodMoat

GoodMoat

GoodMoat is an AI-driven stock valuation tool that breaks away from traditional black-box models. Each valuation figure is directly traced to the original SEC filing, with its source and refresh time clearly noted. It supports full DCF, Reverse DCF (to gauge priced-in growth), and three cross-checked fair-value models for any stock. The X-Ray feature uses AI to deep-dive into 40+ financial metrics, delivering plain-English insights on whether a business has a genuine moat or mere hype. All AI outputs are checked against source filings, ensuring no hallucinated numbers.

Osmosis

Osmosis is a hackathon prototype for a CRM that captures deals from natural team chat instead of forms, presented at the HMD Secure Sales Hackathon 2026.

Q-bit AI pro 2.0

The public page for qbitaipro.com presents itself as a BTC Futures Engine and exposes only a terminal login screen with a demo account. There is no visible feature list, team page, regulatory disclosure, or pricing on the landing page, so this entry sticks to what is verifiable and does not describe capabilities that are not documented.

Pommy AI

Pommy AI is an automation system for founders and marketers that generates, schedules, and optimizes social media posts (reels/shorts) and video ad campaigns. It learns brand voice, designs creatives, targets audiences, and handles cross-platform distribution for growth on autopilot.

Open-source Alternatives

Operit: Open-source Android AI agent connecting models with tools for real tasks

Operit is an open-source Android AI agent primarily written in Kotlin. It connects cloud or local models with system tools, terminals, and browsers to execute real user tasks. As of collection time, it has 5669 GitHub stars and uses an Other license.

OctoBot: Free Open-Source Python Crypto Trading Bot

OctoBot is a free open-source Python crypto trading bot that automates strategies on over 15 exchanges. It includes backtesting, paper trading, and a web UI for easy management. Licensed under GPL-3.0, it has 6146 GitHub stars as of collection time.

Casdoor: Open-source UI-first identity and access management platform

Casdoor is an open-source, UI-first identity and access management platform positioned as a dedicated authentication server. It provides a modern web console for managing users, organizations, applications, and identity providers, with support for OAuth 2.0, OIDC, SAML 2.0, CAS, and LDAP. It includes WebAuthn and passkey support, TOTP-based MFA, biometric login, SCIM 2.0 provisioning, RBAC, and multi-tenant organization models. The stack combines a React frontend with a Go and Beego backend, persisting to MySQL, PostgreSQL, and other databases. The project is licensed under Apache-2.0.

OpenAlice: Local AI Trading Workspace with Git-Style Review Workflows

OpenAlice is a local trading workspace where AI coding agents execute research, portfolio management, and broker orders through Git-style, review-gated workflows. The project is primarily written in TypeScript, licensed under AGPL-3.0, and had 5,201 GitHub stars at the time of collection.

comp: Open-Source AI-Native Compliance Platform

comp is an open-source, AI-native compliance platform that automates SOC 2, ISO 27001, and more. As a self-hosted alternative to Vanta and Drata, it reduces costs and keeps data on your own infrastructure. Built with TypeScript, it offers automated evidence collection, smart policy checks, and risk analysis. Ideal for mid-size teams valuing data sovereignty and customization.

Awesome-LLM4Cybersecurity: Curated Resources for LLM + Security

Awesome-LLM4Cybersecurity is a curated GitHub repository compiling the latest papers, tools, datasets, and frameworks at the intersection of large language models and cybersecurity. Maintained by a community of experts, it claims to have over 1600 stars, making it an essential resource for security researchers and AI developers. The project is primarily written in JavaScript and released under the MIT license.