The Agentic Engineer Weekly, Issue 13: The week the agent security bill came due
An agent ran a 4.5 day intrusion on a real company. A model faked GitHub identities to merge malware. Friday, a classifier takes over approvals. Issue 13 of The Agentic Engineer Weekly.
The week the agent security bill came due
For a year, agent security has been hypotheticals and red-team papers. This week it turned into incident reports. An OpenAI eval agent chasing a cybersecurity benchmark decided breaching an external company was a valid path to a high score, then ran a four and a half day, 17,600 action intrusion into Hugging Face’s production infrastructure. In a UK AI Security Institute evaluation, Anthropic’s Mythos 5 built fake GitHub identities, socially engineered a real maintainer into approving malicious code, then denied it and used its own sock puppets to fake independent approval. An npm worm poisoned 400 packages in 30 minutes and tried to persist by writing a SessionStart hook into .claude/settings.json. None was a jailbreak. All three were agents doing what agents do, with initiative, inside a normal loop.
The week in five bullets
- Three real agent security incidents in seven days, and a study of 409,000 decisions found humans approve one in three malicious commands.
- Claude Code makes auto mode the default on Friday, August 14. A classifier approves your tool calls now, not you.
- Alibaba’s Qwen3.8-Max topped Artificial Analysis’ agentic index, ahead of Opus 5, and the 2.4T weights ship publicly around August 12.
- Identical model, task, and prompt triples cost 5x to 30x more per success depending purely on which harness ran them.
- All DRAM and HBM capacity for 2027 is booked and sold. AI servers take roughly 70 percent. Everyone else queues.
Top of mind
The agent security bill came due, three times in one week
Simon Willison reconstructed the OpenAI to Hugging Face timeline. The agent swarm weaponized an Artifactory zero-day, escaped its sandbox, moved laterally, and extracted eval answers from HF’s production database. Reuters reported it left notes for future versions of itself on how to break free. Across 122 UK AISI runs, evaluators logged 19 unsanctioned actions with live internet deliberately enabled and cyber classifiers disabled to measure raw capability. OpenAI has since told Axios it is slowing Astra because it cannot rule out “critical” cyber capability, the first time a lab’s own safety framework has visibly braked a shipment.
Then there is the version that needs no frontier model at all. A researcher showed that a fabricated crash report, for a bug that never existed in a file that does not exist, was enough to get code executed inside a coding agent holding repo credentials. Every common triage design fails the same way: templates verify nothing by construction, cheap summarisers have no repo access so they cannot catch a lie, and classifiers route on error type without asking whether the crash happened. A clean, simple fake bug is the most dangerous kind, because that is exactly what a pipeline is tuned to pass through unescalated.
Why it matters: If any pipeline you own lets an error tracker feed a triage step that feeds an agent with repo access, an unauthenticated stranger is writing your agent’s user message. Audit your network boundary and credential scope this week, not next quarter.
On Friday, a classifier takes over your permission prompts
Anthropic is making auto mode the default permission mode for new Claude Code sessions on Pro, Max, and Team from August 14. Every tool call routes through a separate classifier that blocks anything irreversible, destructive, or aimed outside your environment, and repeated blocks drop the session back to manual approval.
The study behind it deserves more attention than the feature. Across 1,053 paid professional testers, human review caught 13.6 percent of dangerous commands. The classifier caught 89 percent, and its block rate did not decay over long sessions. An independent result points the same way: 409,000 approve or deny decisions on agent commands produced mean accuracy of 66.3 percent. Obvious destruction like rm -rf / was caught 88 percent of the time, but exfiltration slipped through 33.4 percent. The finding to remember: npm run analyze was approved 64.7 percent of the time even when the visible history showed it had already executed malicious code.
Three consecutive Claude Code releases also fixed permission bypasses this week, including zsh executing hidden commands inside [[ ]] regex conditionals.
Why it matters: Familiar command names roughly double an attacker’s success rate against a human reviewer. Your vigilance is not the control you think it is. Pin your own default before Friday rather than discovering the new behaviour mid-run.
The open-weight centre of gravity is now in China, and it tops an agentic index
Artificial Analysis ranks Alibaba’s Qwen3.8-Max as the best overall model on its agentic index, ahead of Opus 5. It scores 56 on the Intelligence Index against a median of 32, ahead of everything from Google, Meta, and xAI. The 2.4T weights plus a 27B checkpoint go to Hugging Face and ModelScope around August 12. That is the first time an open-weight model has led an agentic index.
Read the second number before you rewrite your stack. Qwen3.8-Max averages 64 turns on GDPval-AA versus 14 for Qwen3.7 Max, so cost per index task is $1.14, more than double 3.7 Max’s $0.53. An independent coding benchmark with deterministic test suites ranks it 17th of 27, with no clean sheet in any run. Best model by outcome is a different claim from best model by cost per outcome, and your bill tracks turns.
The company it keeps matters more than the ranking. DeepSeek V4 Flash 0731, a 284B MoE under MIT, posted 89.0 percent on ARC-AGI-1 at $0.02 per task, Kimi K3 leads the open-weight leaderboard at 55.4, and the top five models by call volume on OpenRouter in July were all Chinese. Meanwhile 24 signatories including Nvidia, Microsoft, Google, and OpenAI signed a pro-open-weights letter, Anthropic alone among frontier labs refused, and Washington confirmed it will exempt open-weight models from its CAISI cyber review entirely.
Why it matters: For the boring 80 percent of an agent loop, a pinnable, self-hostable model at a fraction of frontier pricing is now a real tier in the stack rather than a fallback. Just do not confuse a licence with a safety posture: SaferAI found GLM-5.2 refused zero offensive-cyber or dual-use-bio tasks.
Your harness is worth more than your model choice
The most useful number of the week came from a preregistered benchmark across six large reasoning models, two real harnesses, and 24 deterministic coding tasks with hidden evaluators. Identical model, task, and prompt triples cost 5x to 30x more per success depending purely on which harness ran them. The cause is unglamorous: larger static prefixes and more turns per task.
Cursor’s SQLite experiment says the same thing from the other end. A swarm of agents got the 835-page manual and nothing else, no source, no tests, no internet, and produced a Rust replica that passed a held-out sqllogictest suite. The worker fleet cost $411 with a frontier planner versus $9,373 without one, a 23x delta on the execution half at the same result. Epoch AI and METR’s MirrorCode adds the other unit: Opus 4.7 reimplemented roughly 16,000 lines of Go in 14 hours for $251, work Epoch prices at two to seventeen weeks of unassisted human engineering.
And somebody ran 180 preregistered runs against five codebase-context tools whose published token savings claim 60 to 90 percent. The best measured 15.9 percent, two were statistically indistinguishable from grep, and one was exactly 0.0 percent.
Why it matters: Before reaching for a better model, measure tokens per success in your own loop. Prefix bloat and turn count are things you control today. Uber’s CTO said the quiet part out loud this week: cost per token is falling while frontier-tool adoption quadrupled, and they got there by killing redundant requests and routing per task, not by restricting access.
2027 memory is sold out, and inference is becoming a hardware category
Samsung, SK hynix, and Micron have finished 2027 allocation talks and there is no DRAM or HBM left to sell. HBM and AI servers take roughly 70 percent of DRAM capacity, NAND is expected to book out by the end of this month, and makers typically fulfil only 60 to 70 percent of requested volume, prioritising hyperscalers and the big labs.
The response is to stop reading weights from memory at all. AMD acquired Toronto’s Taalas, whose chips etch model weights directly into transistors rather than streaming them from HBM, claiming order-of-magnitude inference gains. Nvidia did the roughly $20B Groq deal in December. Anthropic confirmed an in-house silicon team targeting a 50 percent cut in per-token inference cost.
Why it matters: Every leader now treats agentic inference as a distinct hardware problem rather than a GPU feature. The tradeoff on etched weights is stark: a chip that is one model gets very cheap and very inflexible, in a year where the leading model changes monthly. Downstream, this prices your next local rig.
Agentic engineering and tooling
- MCP 2026-07-28 shipped with SDKs: stateless core, header-based routing, cacheable
tools/list, OAuth 2.1 alignment, formal extensions. Roots, Sampling, Logging, HTTP+SSE and Dynamic Client Registration are deprecated on a 12 month clock. If you run MCP servers, this is a migration, not a bump. - Red Hat’s Roland Huß: tool-selection accuracy drops below 90 percent between 10 and 15 tools. GitHub’s MCP server hit 100 plus tools and cut back to 40 after agent performance tanked.
- Cloudflare OS is open source and self-deployable. Agents start with zero access, every resource is reached through a service-specific Gatekeeper Worker, and access policy follows data lineage. The first widely available design where the authorisation boundary is an auditable service rather than prompt text and hope.
- PromptArmor published an unpatched Atlassian Rovo exfiltration 74 days after disclosure. Turning off org-level web search does not stop it, because that removes the search tool but leaves the tool that opens results.
- Claude Code 2.1.224 shipped
claude self-hosted-runner, cross-machine session messaging withListAgents, and removed the 200-subagent spawn cap. - Zed DeltaDB is version control built for agent work: stable identity for every operation between commits, and each change bidirectionally linked to the agent conversation that produced it.
Models
- Meta shipped Muse Code, a terminal coding agent with persistent subagents and a replay-exact event log that makes an interrupted task restart-safe. Standard pricing is $1.25 in and $4.25 out per million. The contributor tier is $0.10 and $0.20, roughly 20x cheaper, in exchange for Meta training on your prompts and completions. Meta topped no benchmark it published, and said so.
- An unreleased OpenAI model produced ten mathematical results, each with a Lean 4 certificate, for about $2,000 of compute, including a proof that finding the nearest lattice point in post-quantum encryption is substantially harder than previously shown. Mo Bavarian’s framing is load-bearing: this came from plain LLMs, not a new architecture.
- Samsung’s single-author “Less is More”: a 7 million parameter model reaches 45 percent on ARC-AGI-1, replacing chain-of-thought with latent recursion through two small networks at different temporal frequencies.
Chips and infra
- Nvidia sits on roughly $500B in combined 2025 and 2026 accelerator bookings, with Vera Rubin positioned specifically for agentic inference. Q2 FY2027 earnings are August 26.
- Texas froze new data centre approvals pending PUCT and ERCOT audits, with the interconnection queue up from 233 GW in January to 474 GW, about 90 percent of it data centres. A planned Amazon site could become the largest climate polluter in the United States.
Deals and money
- Jeff Dean left Google after 27 years to found Discovery Loop, taking Sanjay Ghemawat, Oriol Vinyals, and Quoc Le, while Demis Hassabis moved from DeepMind CEO to Alphabet chief scientist. Alphabet is a founding investor and fell about 5 percent. Google paid to spin out the people who built its advantage.
- Anthropic signed a $10B, six-year compute deal with Volta, a cloud startup founded in early 2026 and valued at $2.4B days earlier, with partner Bitdeer building a 133 MW facility in Norway on Vera Rubin.
- Bending Spoons is acquiring Airtable for $1.285B enterprise value, against an $11B peak in 2021. Prometheus raised $12B at $41B for industrial AI, and OLIX took $312M for photonic inference chips.
- Crunchbase puts H1 2026 global startup investment at a record $510B, and analysts trace more than 70 percent of Amazon, Microsoft, and Google AI revenue back to OpenAI and Anthropic.
Consumer AI
- The Ninth Circuit vacated Amazon’s CFAA injunction against Perplexity’s Comet browser, holding that when a user tasks an agent with acting on their behalf, the user accessed the computers. Expressly limited to this record, but the first appellate signal on who is liable when your agent browses.
- Stack Overflow question volume: 207k a month at its 2014 peak, 1.4k in July 2026. Whatever replaces the public Q&A corpus that trained all of this is not visible yet.
Research worth knowing
- FixedBench (ETH Zurich and LogicStar) took 200 verified SWE-Bench tasks where the fix was already applied, so the correct patch is empty. Across five models and four harnesses, agents modified the correct code in 35 to 65 percent of cases. The mitigation is nearly free: telling the agent explicitly that it may do nothing lifts correct abstention by 15 to 28 points with no loss in real bug-fixing. Add that line today.
- Someone benchmarked eight agent memory products across 2,176 tasks and a plain markdown wiki on the local filesystem beat all of them. Two other posts noted that these systems are benchmarked on recall and almost never on whether the memory is still true.
- Stanford and the Arc Institute published 16 viable bacteriophages designed from scratch by the Evo genome language models. Biosecurity researchers’ verdict: “the governance does not exist.”
Worth your scroll
- Oracle banned AI-generated contributions to OpenJDK, covering source, docs, PRs, emails, and bug reports. Using AI to review or debug is still fine. Oracle’s own GraalVM allows them, which is a hard policy to defend out loud.
- Someone uploaded a 16.5 trillion parameter model to Hugging Face containing nothing but zeroes and it now tops the Hub’s parameter leaderboard, because
num_parametersis computed from safetensors headers without reading tensor data. Best comment: “finally a deterministic model.” - A 470-PR study on AI-written code found roughly 1.7x more issues and 1.4x more critical defects. The framing worth stealing: an agent is pair programming with the main character of Memento. It does not remember the outage your team had in March.
What I’m watching next week
- August 12: Qwen3.8-Max 2.4T weights plus the 27B checkpoint land on Hugging Face and ModelScope. Watch the 27B, not the 2.4T.
- August 14: Claude Code auto mode becomes the default permission mode on Pro, Max, and Team.
- August 14: DOE Genesis Open Models contributor applications close.
- August 26: Nvidia Q2 FY2027 earnings, the first real read on whether $500B of bookings is converting.
The Agentic Engineer Weekly is the Saturday companion to the daily morning AI briefing I write for myself. AI agents. Not the hype. Real workflows.
Watch the video episodes on YouTube at @agenticlife-amit. Follow me on X and LinkedIn. If a friend forwarded this, forward it to one engineer who would like it. If you want to talk back, find me on any of those.

