FrontierHarness Eval (opens in a new tab)
Runta
FrontierHarness holds the model and runtime constant while comparing different coding-agent harnesses across quality, cost and latency.
Issue #004 · September 13, 2026
The model is no longer the whole system. This week, the interesting failures — and some of the biggest performance differences — are showing up in the harness, context assembly, tool boundaries, evaluation layer and execution environment.
Runta
FrontierHarness holds the model and runtime constant while comparing different coding-agent harnesses across quality, cost and latency.
ContextOS
An independent recomputation examines FrontierHarness's headline economics and highlights how denominator choices and missing cost data affect interpretation.
arxiv.org
Researchers systematically analyze how 12 real-world coding-agent harnesses assemble context and identify two new attack classes: MessageRole Context Privilege Escalation and Cross-Scope Context Privilege Escalation.
GitHub Security Advisory
When MCP support is enabled, affected Chainlit versions allowed unauthenticated clients to make the server issue arbitrary outbound requests, including requests toward internal services and cloud metadata endpoints.
Open-source project
Tansive experiments with declarative access policy around agent tools, session-level data constraints and authenticated MCP exposure.
Modal engineering
Modal explains the scheduler and control-plane redesign behind creating and running extremely large numbers of isolated sandboxes for agents and RL workloads.
arXiv
AgentProp-Bench analyzes 14,750 execution traces across 13 agents to study evaluator reliability, error propagation and tool-execution hallucination.
Get the next one straight to your inbox.