AI DevList

Issue #002 · August 30, 2026

AI DevList #002

This week, the interesting shift isn’t another model release. It’s what happens around the model: how agents are contained, evaluated, given memory, connected to tools, and governed once they start operating real systems.

That’s where agent engineering starts to look a lot more like systems engineering.

Agentic AI

AI Security & Red Teaming

VMs won't contain cyber-capable agents (opens in a new tab)

Trail of Bits

Trail of Bits gave a cyber-capable agent a VM-escape challenge. It used known vulnerabilities, combined bugs that were not marked as security issues, and eventually constructed an exploit chain involving several zero-days. Trail of Bits tested whether a modern cyber-capable agent could escape a conventional QEMU/KVM sandbox. It succeeded through multiple paths, eventually chaining newly discovered vulnerabilities.

LLM Engineering

Tools & Frameworks

Effective Patterns for Advanced MCP Usage (opens in a new tab)

PulseMCP / O’Reilly

The article moves beyond single MCP-server demos and looks at compositions where multiple clients and servers work across applications and workflows. The real value of MCP starts when workflows cross application boundaries. At that point, configuration, authentication, tool discovery and governance become much more interesting than simply exposing one API to one agent.

DeepSeek Harness (opens in a new tab)

DeepSeek / GitHub

DeepSeek Harness is an MIT-licensed agent harness built around a plugin-based architecture where runtime capabilities can be composed and replaced. It is explicitly still a developer preview. The harness layer is becoming as important as the underlying model: tools, state, sessions, permissions and execution loops increasingly live outside the LLM itself.

Research & Papers

Production & Evaluation

Find Evil! — autonomous incident-response agents tested by practitioners (opens in a new tab)

SANS

This one is excellent. 90 practicing incident responders ran 1,775 evaluations against 123 agent harnesses, including adversarial testing. The top five implementations are open source. Most agent benchmarks are constructed by AI researchers. Here, actual incident responders tested whether autonomous agents could investigate compromised systems while resisting destructive or misleading behavior.

Tutorials & Deep Dives

Enjoyed this issue?

Get the next one straight to your inbox.