Patterns and problems in emerging multiagent systems (opens in a new tab)
Anthropic
Anthropic explores what happens when multiple frontier agents interact in shared environments, including cooperation, competition and unexpected systemic behaviors.
Issue #001 · August 23, 2026
Multi-agent systems, MCP, agent skills, observability and production AI.
A focused selection of what mattered this week for engineers building AI systems.
Anthropic
Anthropic explores what happens when multiple frontier agents interact in shared environments, including cooperation, competition and unexpected systemic behaviors.
LangChain
LangChain explores generating small orchestration programs that coordinate subagents instead of relying exclusively on sequential model tool calls. The approach targets context isolation, parallel execution, branching, and more efficient orchestration.
Pillar Security
Pillar Security uncovered an active MCP supply-chain campaign where a seemingly harmless server changes its tool metadata after a few normal calls, then attempts to steer the connected agent toward sensitive credentials and local configuration.
AgentAuditKit / independent open-source research
An open-source security study scanned 571 real public MCP configurations and found recurring weaknesses around authentication, package pinning, secret exposure, shell execution and supply-chain risk.
Hugging Face
Hugging Face examines open-model activity across the first seven months of 2026, including releases, downloads, licensing patterns, derivatives and the increasing role of agents on the Hub.
Cloudflare
The MCP specification introduces a rewritten stateless core alongside updated TypeScript, Python, Go and C# SDKs.
Microsoft
Structured outputs, web search, web fetch, MCP connector and tool search are now available for Claude deployments hosted on Azure.
arXiv
A new study looks at when procedural skills actually improve agent behavior and where skill-augmented agents still break across realistic execution environments.
Arize / Uber
Uber's experience shows why offline agent evaluation isn't enough: production traces reveal user behaviors, tool disagreements and failure modes that pre-launch datasets don't necessarily capture.
Cloudflare
Cloudflare open-sourced the agent workspace platform it has been running internally, combining organizational context, skills, isolated runtimes and governed access to internal resources.
Google Cloud
Google describes its process for managing skills as maintained engineering artifacts, including validation, CI and recurring evaluation rather than treating them as static prompt files.
LlamaIndex
ExtractBench evaluates 14 extraction systems across 370 enterprise documents and 4,869 pages, measuring factors including completeness, grounding and cost. Its dataset and evaluation harness are publicly reproducible.
Get the next one straight to your inbox.