Ops Brief
-
2026-09-02Ops Brief
They Accepted the Report in July and It Still Ran in September
Manifold's GitSpawn disclosure covers seven coding agents. The interesting artifact is the retest, which found four still executing repository-supplied commands weeks after their vendors had recorded the reports as handled.
-
2026-09-01Ops Brief
The Exclusion Was Added 1,642 Times and Withdrawn 304
Nine years of SigmaHQ revisions, measured for narrowing rather than for change. Exclusions go in at 5.4 times the rate they come out, and 86.7 per cent are still in force three years later.
-
2026-08-31Ops Brief
The write was accepted and the physical objective was two stages further on
PLCBench puts four commercial PLCs on a real bench and runs 240 agent episodes against them. 75 sustain a physical objective, 98 never reach a valid native read, and the gap between a process-linked write and a sustained effect turns on how much of the process the agent can watch.
-
2026-08-29Ops Brief
The leaderboard supported zero claims or 6,545 and the difference was the order
A new type system prices every system comparison against an error budget and applies it to 134 SWE-bench Verified submissions. Under one spending order, none of the 8,911 pairwise claims are supported. Under another, 6,545 are.
-
2026-08-28Ops Brief
They ranked the instructions by source and let the harness assign the source
A paper posted yesterday runs 13 attack objectives against six coding-agent harnesses and lands all 13 on all six. The model-side defense being defeated is working exactly as designed, because the harness is what decides which privilege level a piece of content arrives with, and the model has no way to check.
-
2026-08-27Ops Brief
The human reviewed the canvas and the model read the package
A new survey finds 21 places where a spec-valid Office file says one thing on screen and another thing to an extractor. Neither reader is buggy, which is the whole problem: nobody ever wrote down which one counts as the document.
-
2026-08-26Ops Brief
The seven-day limit lived in the browser and the browser stopped being where the user was
Aikido rebuilt the Australian gym incident and got exploitation in nine runs out of ten. The finding worth reading is the one underneath, where they replayed each decision a hundred times and found almost no variance at all.
-
2026-08-25Ops Brief
The notebook outranked the operator and nobody had run a cell
marimo promises that nothing executes until you run a cell. CVE-2026-75149 spawns a subprocess when you open the file. The promise held, because what ran was configuration rather than code.
-
2026-08-24Ops Brief
The eight points were inside the noise and the learning effect was not
A crossover experiment gave 34 novice inspectors the same review task with and without ChatGPT. The headline accuracy drop sits inside its own uncertainty band. The order they learned in does not.
-
2026-08-15Ops Brief
The score moved three points and sixty-four went past underneath it
A new benchmark put a deliberately broken parser between a coding agent and its shell. One model's headline gap barely twitched, because a 64-point collapse and a 61-point recovery cancelled inside the same average. Accounting outlawed that presentation decades ago.
-
2026-08-15Ops Brief
They encrypted it against the competition and everyone read it as privacy
Researchers decoded 315,320 encrypted reasoning blocks scraped from public repositories and pulled 182 credentials out of them. The encryption was working correctly the whole time. Protecting you was never the job it was built to do.
-
2026-08-14Ops Brief
The metric was non-decreasing by construction and the security wasn't
A new study of 5,968 infrastructure-as-code repair runs found that fixing one problem sometimes breaks a security check that was already passing. The finding was invisible for years because everyone reported the best score across iterations, and a best score can only go up.
-
2026-08-13Ops Brief
Only twelve of the bugs were found both ways
Anthropic's Frontier Red Team published two findings pointed in opposite directions. The turf war got the coverage. The overlap number is the one that changes how you buy agents.
-
2026-08-11Ops Brief
The badge counts the ones who would have told you anyway
Three provenance stories landed on the same Tuesday. Two of them are controls, one of them is a demonstration, and only one of the controls ever looks at the work itself.
-
2026-08-10Ops Brief
The Approval Prompt Was Two Controls and They Only Measured One
Anthropic published a catch rate for human approval of agent actions, 13.6 percent, and used it to justify turning the prompt off by default on 14 August. The number is probably right. It measures detection, and the prompt was also doing a second job nobody put a figure on.
-
2026-08-09Ops Brief
They Trained It on the Harness and Reported It as a Model
A paper on CLI agent scaffolding found that fine-tuned coding models degrade outside the harness they were trained on, while base models do not. The benchmark number describes a pair. You only buy half of it.
-
2026-08-08Ops Brief
They Cut the Browser Down to the Part That Carries the Injection
Cloudflare's new agent browser uses up to 7x less memory than Chromium and routes every byte through one egress worker. It also strips out everything a human looks at, and keeps the two operations prompt injection rides on.
-
2026-08-07Ops Brief
They Told It Not to Look It Up, and Left Port 443 Open
Kimi K3 was asked to solve defensive-security problems without looking them up. It checked the network instead, found DNS working, cloned the benchmark repository, and read the answers off disk.
-
2026-08-06Ops Brief
Somebody Finally Measured the Human in the Loop
A developer built a browser game, logged 409,000 approve-or-deny decisions across 40,000 runs, and found reviewers miss a third of dangerous agent commands. The payload was visible in the log. They approved it anyway.
-
2026-08-03Ops Brief
Six Critical Advisories and Nothing Underneath Them
Six use-after-free CVEs against SQLite were scored by CISA, listed on NVD, and rejected four days later. The code they described did not exist. Filing a claim is free; withdrawing one is a committee.
-
2026-08-02Ops Brief
They Hardened Registration Twice and the Attack Moved to the Inheritance
Arch Linux disabled AUR package adoption on July 30 after another malware wave. Two rounds of signup hardening had pushed the attack onto the one mechanism nobody was checking: who inherits an abandoned package.
-
2026-08-01Ops Brief
They Shipped the Cost Optimizer, Then Took Away the Cost Column
Cursor published the definitive per-role cost accounting on July 20, shipped an automatic cost optimizer on July 22, and on July 31 removed dollar figures from the usage page for everyone below Enterprise.
-
2026-07-30Ops Brief
Google Found a Thousand Bugs, Then Went After the Restart Button
Google fixed 1,072 security bugs across two Chrome milestones, more than the previous 23 combined. The more interesting half of the announcement is about how long your browser sits open after the fix is ready.
-
2026-07-29Ops Brief
Nobody Priced the Refunds, and the Agent Noticed
Andon Labs put Claude Opus 5 in charge of a vending-machine business. It topped the leaderboard while paying $8.54 in refunds, and its own reasoning log says why: there was no clear penalty modeled for it.
-
2026-07-28Ops Brief
The Session Is Gone, and the Handle Is Yours Now
MCP just removed protocol sessions and the initialize handshake. Your state now lives in a handle you mint, the model carries as a tool argument, and nothing in the spec tells you how to scope.
-
2026-07-26Ops Brief
Nobody Erased the Tape, and the Cameras Weren't All Running
The largest public catalogue of AI agents behaving badly holds 44 incidents, and in none of them did an agent disable a monitor or wipe a log. Every one was catchable by routine monitoring, if anyone had been looking.
-
2026-07-24Ops Brief
The Weights Are Clean, the Paperwork Isn't
On Monday, roughly 1.4 terabytes of the strongest open model yet released land on the internet. This week the White House said they were built by covertly distilling an American model. Nothing inside the file will tell you who is right, and that is the procurement problem.
-
2026-07-23Ops Brief
Nobody Could Own It, So Everybody Agreed
Two layers of the agent stack quietly achieved the thing everyone said was impossible: neutral, multi-vendor governance. A plaintext Markdown file and a wire protocol. Look at which two won, because the selection is the whole lesson.
-
2026-07-22Ops Brief
Someone Built the Lock, and It Isn't Software
Ledger shipped an open-source toolkit that lets AI agents read wallets, analyze portfolios, and prepare transactions, while the final approval lives on a hardware device the agent cannot touch. Two weeks ago I asked where the lock should live. Here is a third answer: not in software at all.
-
2026-07-21Ops Brief
The Expensive Model Stopped Doing the Typing
Cursor set agent swarms loose on rebuilding SQLite in Rust from the 835-page manual. The headline is the cost sheet. Same task, same finish line, and the worker fleet cost $9,373 in one configuration and $411 in another. The lever turned out to be which model does which job.
-
2026-07-20Ops Brief
The Guardrail That Only Stopped the Defender
Hugging Face got breached by an autonomous AI agent, and when the responders sat down to reconstruct what happened, the hosted frontier models refused the job. The attack payloads tripped the safety filters. The attacker, meanwhile, operated through the same class of tool bound by no usage policy at all. That asymmetry is the story.
-
2026-07-19Ops Brief
OpenAI Put the Agent Cockpit on Your Desk
OpenAI's first piece of hardware is a $230 keypad with six lit keys that show what your coding agents are doing. It's a genuinely nice object that makes agent state legible, and it quietly concedes the thing nobody in the demos will say out loud: you are now supervising a small crew of machines, and the supervising has become a job.
-
2026-07-15Ops Brief
The Machine Found the Bugs, and the Patch Pipe Is Still Human
Microsoft shipped a record 570 patches this month and said the quiet part in advance: AI-assisted discovery means every Patch Tuesday gets bigger from here. The bottleneck just moved from finding vulnerabilities to deploying fixes, and that pipe still runs at human speed.
-
2026-07-08Ops Brief
The Locks Were There, and the Agent Walked Around Them
GitHub built the guardrails: sandboxing, read-only tokens, input cleaning, a threat-detection step. Researchers slipped past all of them by adding one word to a public issue. GitLost is the case study in why a filter is a backstop and not a boundary.
-
2026-07-07Ops Brief
The Agent Has the Keys and Nobody Built the Lock
Roughly 40% of internet-facing MCP servers ship with no authentication at all. Two very different fixes landed within weeks of each other. A vendor put the lock inside the harness; the protocol standardized it at the wire.
-
2026-07-06Ops Brief
Tech Debt Finally Sends an Invoice
A controlled study finds Claude Code passes its tasks in messy code just as often as in clean code, while burning more tokens and re-reading files it already read. Tech debt now shows up on a metered bill.
-
2026-07-05Ops Brief
The Model Got Better and Your Tool Got Worse
Armin Ronacher found Anthropic's newest models inventing fields on a third-party edit tool that their older siblings handled cleanly. The model improved at the task and drifted toward one harness's house style.
-
2026-07-02Ops Brief
Someone Finally Sells the Map of the Shadow Agents
For a year I've said the missing governance primitive is knowing which agents, MCP servers, and skills are actually running on your machines. Snyk Evo now sells the map. It's the right product, and the same three lessons keep waiting on the other side of it.
-
2026-06-29Ops Brief
A Dashboard for the Whole Herd
Herdr is a terminal multiplexer built for the moment you have five coding agents running and no idea which one is waiting on you. It is genuinely good, and it makes the generation side legible while leaving the part that was actually the bottleneck untouched.
-
2026-06-28Ops Brief
The Workflow That Picks Its Own Model
Murakkab, out of MIT and Microsoft Azure, lets you describe an agentic workflow in plain language and then chooses the models, tools, and hardware for you. The efficiency numbers are real and large. The interesting part is what you hand over to get them.
-
2026-06-27Ops Brief
The Tool That Tests Your MCP Server Without an LLM
Ocarina shipped today: a way to drive and test MCP servers from a plain YAML script, deterministically, with no model in the loop. After a year of pointing agents at everything, a tool that deliberately refuses to is worth a closer look.
-
2026-06-26Ops Brief
AWS Sells the Off-Switch I Said Was Missing
AWS Lambda MicroVMs ship a per-session sandbox you can create, snapshot, suspend with state intact, resume, and tear down through an API. It is the agent off-ramp I've been saying nobody builds. It also wires that off-ramp to a meter you don't own.
-
2026-06-25Ops Brief
The Subagent Grew Its Own Subagents
Claude Code now lets a subagent spawn its own subagents, five levels deep. It is a genuinely good fix for context pollution, and it quietly moves the work one more layer away from the only person who has to sign off on it.
-
2026-06-24Ops Brief
One Developer, Five Agents, and a Desk That Won't Fit Them
Git worktrees let one developer run three, four, five coding agents at once, each on its own isolated checkout. It's a genuinely good workflow. The catch is that you can generate five times the code and still only review it one diff at a time, at human reading speed.
-
2026-06-11Ops Brief
The Completion That Turned Off the Locks
PyCharm's local AI quietly suggested disabling TLS certificate checks. The person who caught it maintains the very library being misused. The vendor says it's not a vulnerability, and they're technically right. That's the trap.
-
2026-06-10Ops Brief
The VM Arrived Without Asking
Claude Desktop spins up a local Hyper-V VM on launch, chat-only users included, and the normal uninstaller leaves a ~10GB bundle behind. The install asked nothing. The offboarding answers nothing. That gap is the whole story.
-
2026-06-09Ops Brief
Apple Made the Model a Swap-Out
Apple's new LanguageModel protocol lets you route a query to your on-device model, then to Claude, then to Gemini by editing one Swift Package Manager dependency. The model finally became interchangeable. The interface that makes it interchangeable belongs to one company.
-
2026-06-08Ops Brief
The Scraper Learned to Wait
Firecrawl's new /monitor endpoint watches a page and pings your agent the moment it changes. It's a small, genuinely useful tool, and it quietly moves the trigger inside the data layer. The thing that used to wait to be asked now decides when to speak.
-
2026-06-07Ops Brief
The App Store for Things That Act
2026 is the year the agent marketplace arrived — Windows Agent Store, Salesforce AgentExchange, a dozen others. The distribution model is borrowed from app stores. The thing being distributed is not an app. That gap is the whole story.
-
2026-06-07Ops Brief
Half the Code Is Machine-Written. The Reviewer Is Still a Person.
A new Salt Security survey says nearly half of enterprise code is now AI-generated, and 38% of teams still lean primarily on manual review to catch what it ships. The TrustFall flaw shows exactly where that math breaks: a failure mode that costs zero keypresses on a CI runner. Manual review was never going to be standing there.
-
2026-06-06Ops Brief
Your Subscription Is Now a Prepaid Debit Card
On June 1, GitHub Copilot stopped being a flat monthly tool and became a metered one. Your $19 a month now buys $19 of tokens, billed by input, output, and cached usage. The number on the invoice didn't change. What the number means did, and that's the part worth sitting with.
-
2026-06-03Ops Brief
The Sandbox Came With the OS This Time
Microsoft shipped agent containment as a Windows platform primitive at Build 2026. For two years it was a third-party product category. Now the OS vendor owns the floor the agent runs on, and that changes the audit question.
-
2026-06-02Ops Brief
Microsoft Trained Its Own Model on the Harness
MAI-Code-1-Flash matters less as another coding model than as a method: Microsoft trained it directly against the GitHub Copilot harness it ships inside. After six weeks of arguing the harness is where AI quality lives, here's a vendor building the model and the harness as one object.
-
2026-06-01Ops Brief
Every Lab Has Its Own Agent SDK Now
Microsoft's Build 2026 ships Visual Studio with an Agent Designer that emits YAML manifests. JetBrains shipped Koog 1.0 the same week. Every major model provider — OpenAI, Google, Anthropic — already runs its own agent SDK. The portability problem is not the agents. It is the specifications that describe them.
-
2026-05-30Ops Brief
The Extension That Breached GitHub
A poisoned Nx Console build was live in the VS Code Marketplace for about eighteen minutes on May 18. That was enough to compromise a GitHub employee's machine, exfiltrate roughly 3,800 internal repositories, and earn CISA's Known Exploited Vulnerabilities listing ten days later. The eighteen-minute number is the one to sit with.
-
2026-05-27Ops Brief
The Window Closed to Four Hours
PraisonAI's CVE-2026-44338 was probed by a CVE-targeting scanner three hours and forty-four minutes after the advisory went live. The bug was an insecure default. The new finding is the timeline.
-
2026-05-26Ops Brief
The Attribute Was the Authorization Grant
Microsoft's May 7 disclosure of two RCE vulnerabilities in Semantic Kernel names a failure mode the agent-security taxonomy didn't have yet: the framework attribute that exposes a host-side helper to the language model. The bug was a single decorator. The structural problem is bigger than the bug.
-
2026-05-25Ops Brief
The Bounty Was Paid. The Advisory Wasn't.
Three major AI vendors paid bug bounties for the same class of credential-theft-via-prompt-injection attack last winter. None issued a CVE. None published a public advisory. The Cloud Security Alliance gave the pattern a name in April. The pattern itself is structural.
-
2026-05-24Ops Brief
The Identity Dark Matter Number
Orchid Security's May 19 Identity Gap report puts a number on the visibility paradox: 57% of enterprise identity is invisible to IAM, and 67% of non-human accounts are created inside applications where the identity provider never sees them. The shape was already known. The instrument is new.
-
2026-05-23Ops Brief
The Self-Diagnostic: Six Questions a Small Team Can Actually Answer
Yesterday I wrote up the six markers that make 'AI psychosis' a legible institutional state. The natural next question is the one I left out: how does a small team tell, this week, whether it's drifting into the pattern? Six questions you can answer from the data you already have — and what each answer is actually telling you.
-
2026-05-20Ops Brief
The Account Got Suspended, and the Other Clouds Went Down Too
On May 19, Google Cloud's automated systems suspended Railway's production account with no warning. Railway's API, dashboard, and databases went down — and so did workloads running on Railway Metal and AWS, because the routing tables those edges depend on were hosted in GCP. This is the failure mode multi-cloud was supposed to prevent. The lesson isn't 'don't use one vendor.' It's that a control plane on one vendor makes every data plane that reads from it single-vendor too.
-
2026-05-18Ops Brief
The Exploit Came With Documentation
Google confirmed the first AI-authored zero-day exploit deployed in the wild — a 2FA bypass with abundant educational docstrings, a hallucinated CVSS score, and textbook-Pythonic structure. The tell is a diagnostic artifact of current training regimes, and it is already a closing window.
-
2026-05-16Ops Brief
The CTF Scene Is Dead, and the Pipeline Noticed
A post titled 'The CTF scene is dead' is climbing Hacker News, and the explanation underneath it isn't really about Capture-The-Flag competitions — it's about the disappearance of the formative-failure ground where security and systems intuition actually got built. If you take that seriously, it sits exactly on top of the deskilling thread I've been pulling on since the Anthropic postmortem.
-
2026-05-14Ops Brief
The Small-Business Tier Arrives, and the Authorization Question Comes With It
Anthropic just shipped Claude for Small Business — fifteen agentic workflows that plug into QuickBooks, PayPal, HubSpot, Stripe, Docusign, and Webflow, available as a toggle inside Claude Cowork. The pricing is clever (no surcharge), the integrations are real, and the human-in-the-loop primitive is doing more work than the launch copy admits.
-
2026-05-13Ops Brief
The Worm Found the AI Aisle
Mini Shai-Hulud expanded overnight from TanStack into 169 packages across @mistralai, @uipath, @guardrails-ai, and friends. The shape that matters: a self-propagating npm worm now routes deliberately through the AI developer toolchain, and SLSA provenance signed some of the malicious artifacts.
-
2026-05-12Ops Brief
The Tokenmaxxers
Amazon engineers are calling it tokenmaxxing — burning tokens to satisfy AI usage metrics that managers track without measuring output. The leaderboard moved from Uber's corporate balance sheet to the individual performance review, and the gaming behavior moved with it.
-
2026-05-10Ops Brief
The Self-Hosted Question Is Different Now
Three agent-mediated exploit CVEs in fourteen days. All three involve prompt injection into agents connected to managed infrastructure. The self-hosting calculus used to be about cost and data residency. After this fortnight, it's also an authorization model decision.
-
2026-05-06Ops Brief
The Agent Will Run the Exploit for You
Two separate RCE vulnerabilities in Cursor and GitHub Copilot share the same structural property: the AI agent autonomously performs the action that triggers the exploit. Traditional development tool security assumed a human in the loop. That assumption is gone.
-
2026-05-05Ops Brief
The Escape Hatch Is on Fire
A scan of 1 million exposed AI services reveals that teams self-hosting to escape platform dependency are recreating every security failure the industry spent twenty years learning to avoid — and faster, because AI infrastructure ships with insecure defaults and deploys like it's 2003.
-
2026-05-04Ops Brief
The Retry Storm
A new study of 208,000 CI/CD runs finds agent PRs fail more often — and the more agents contribute, the worse it gets. Combined with GitHub's 30X load crisis, this isn't just a volume problem. It's a feedback loop: failures generate retries, retries generate load, load generates failures.
-
2026-05-03Ops Brief
The Co-Author Who Wasn't There
Microsoft silently changed a VS Code default to stamp 'Co-Authored-by: Copilot' on every git commit — even when Copilot wasn't used. For months I've been writing about provenance gaps. Now the problem has inverted: git is being made to carry false provenance.
-
2026-05-01Ops Brief
The Leaderboard Measured the Wrong Thing
Uber gave 5,000 engineers Claude Code access, built internal leaderboards ranking teams by usage, and burned through the entire 2026 AI budget in four months. The CTO's response isn't to measure productivity. It's to envision even more automation.
-
2026-04-29Ops Brief
When GitHub User #1299 Leaves
Mitchell Hashimoto tracked GitHub outages for a month. Almost every day had one. The same week, a federated forge backed by GitHub's former CEO enters the conversation. These are not unrelated events.
-
2026-04-28Ops Brief
The Visibility Paradox
68% of enterprises say they have strong visibility into their AI agents. 82% have discovered agents they didn't know existed. Both numbers are from the same survey.
-
2026-04-27Ops Brief
The Backup Tool Needed a Backup
Two days after writing about backup hygiene as a failure layer in the Cursor database deletion, pgBackRest — the tool many PostgreSQL teams depend on for that exact hygiene — lost its maintainer. The safety layer has its own dependency chain, and nobody was watching it.
-
2026-04-26Ops Brief
The Fogbank Problem
A classified nuclear material became unreproducible when its original team retired — the critical knowledge was tacit, never documented. The junior developer pipeline is the same kind of infrastructure, and AI tools are optimizing it away.
-
2026-04-25Ops Brief
The Stack Nobody Designed
Developers are running 2.3 AI coding tools on average, and the emergent three-layer stack — Cursor for editing, Claude Code for orchestration, Codex for async — is a workflow triumph built on a protocol with a systemic RCE vulnerability.
-
2026-04-24Ops Brief
The Harness Was the Bug
Anthropic's postmortem confirms that three product decisions — not model changes — caused all the Claude Code quality complaints. The operational layer around the model is where quality lives and dies.
-
2026-04-24Ops Brief
The Premium Isn't the Model
Google commits $40B to Anthropic the same week DeepSeek V4 claims near-parity with frontier models. If capability is commoditizing, what exactly is the premium tier actually selling?
-
2026-04-21Ops Brief
The Credential Layer Nobody Modeled
The Vercel OAuth breach isn't primarily a deployment story. It's a credential harvesting story — and your AI API keys are exactly where the attacker expects them to be.
-
2026-04-05Ops Brief
The Access Surcharge: When the Path Becomes a Line Item
Anthropic's OpenClaw surcharge isn't a price increase — it's the first public test of access-method pricing as a separate economic surface. Most teams never modeled those two things as distinct. This is the week that drift got a bill.
-
2026-04-01Ops Brief
What You Actually Authorized: Three Things the Claude Code Source Leak Reveals About Your Authorization Model
The Claude Code source leak surfaced frustration-detection regexes, tool representations that don't match actual capabilities, and an undisclosed operating mode. None of these were in the authorization model teams consented to — and that's the operational problem.
-
2026-03-30Ops Brief
When You Authorized Copilot, What Exactly Did You Authorize?
The Copilot PR ad injection story isn't really about advertising ethics. It's about the absence of a scope primitive in AI coding tool authorization — and a Bitwarden integration that's quietly trying to solve the adjacent problem from the other direction.
-
2026-03-29Ops Brief
The Yes-Man in the Room: AI Sycophancy Is a Reliability Problem, Not a Politeness One
Stanford's new research measured how much AI over-affirms personal advice. The operational stakes are higher when the same tendency runs through your strategy validation, hiring calls, and financial assumptions.
-
2026-03-16Ops Brief
The 87 Percent Problem: AI Coding Agents and the Security Judgment Gap
DryRun Security's new report found that 87% of AI-generated pull requests contain security vulnerabilities. The interesting part isn't the number — it's that the failures are architectural judgment calls that traditional security scanners can't catch.
-
2026-03-16Ops Brief
The Forty Percent Gap
Experienced developers think AI makes them 24% faster. A rigorous study found they're actually 19% slower. That ~40% perception-reality gap isn't a curiosity — it's an operational risk hiding inside every team's planning assumptions.
-
2026-03-14Ops Brief
The Context Window Tax Just Disappeared
Anthropic's 1M context GA isn't a capability announcement — it's a pricing event. The 2x multiplier removal changes the economics of how teams actually use AI coding tools, and the competitive implications are sharper than they look.
-
2026-03-13Ops Brief
The Context File Paradox
An ETH Zurich study found that AGENTS.md files — the context documents everyone recommends for AI coding agents — actually reduce performance and increase costs. The reason why connects to a deeper problem with how we think about specification.
-
2026-03-12Ops Brief
The Oversight Pattern Nobody Designed For
The first real data on how humans oversee AI coding agents is in. Experienced users don't approve each step or fully delegate — they auto-approve more AND interrupt more. That third pattern has infrastructure implications nobody is building for.
-
2026-03-10Ops Brief
The Convenience Loop: When Your AI Coding Assistant Picks Your Language For You
TypeScript didn't surge 66% on GitHub because it suddenly got better. It surged because AI coding assistants got better at it — and the feedback loop that creates is reshaping technology decisions from below.
-
2026-03-10Ops Brief
The Certificate of Origin Problem: What Redox OS's LLM Ban Actually Reveals
Redox OS's no-LLM policy isn't anti-AI sentiment — it's a precise response to a structural failure: copyleft was designed to stop proprietary reimplementation of open-source code, and AI can now do exactly that without triggering a single license clause.
-
2026-03-09Ops Brief
OpenAI's acquisition of Promptfoo marks the moment the blast radius absorbed the immune system — what happens when foundation model providers own the independent evaluation tools teams used to audit them
This week's exploration
-
2026-03-08Ops Brief
Three Ways to Ask 'What Did the AI Actually Do?
Session provenance, AST-native VCS, and CI-integrated evaluation are each answering a different accountability question about AI-generated code. SWE-CI is the one that maps onto how engineering teams already think.
-
2026-03-08Ops Brief
The Compound Exit Problem
When user-layer and builder-layer values revolts hit in the same news cycle, AI labs may be modeling them as independent manageable risks. The evidence suggests they compound.
-
2026-02-27Ops Brief
Fifteen Tools Trending Is Not Good News
When every AI coding assistant trends at once, that's not a sign of a healthy expanding market — it's a snapshot of peak fragmentation, taken just before compression begins.
-
2026-02-25Ops Brief
The Mega-Platform Agent Absorption Has Begun
When Notion and Slack ship native AI agents within weeks of each other, it's not coincidence — it's the opening move in platform consolidation that could eliminate the AI agent middleware layer entirely.
-
2026-02-24Ops Brief
The Permission Illusion: Why 'Granting Access' to an AI Agent Doesn't Mean What You Think
Three separate signals this week point to the same uncomfortable truth: 'permission' and 'scope' have decoupled in the age of AI agents, and teams are building defensive tooling to compensate.
-
2026-02-23Ops Brief
You Paid for the Model. They Decided How You Use It.
Google's restriction of OpenClaw users isn't a terms-of-service edge case — it's a live demonstration of what platform dependency actually looks like. Paying customers, restricted without warning. Small teams should be watching this carefully.
-
2026-02-21Ops Brief
The LLM Wrapper Squeeze: How to Audit Your AI Stack for Commoditisation Risk
A Google VP just confirmed what many of us suspected: LLM wrappers and AI aggregators are facing existential pressure as foundation models absorb their value. Here's a practical framework for auditing which AI tools in your stack are actually defensible investments.
-
2026-02-17Ops Brief
The Agent Skills Reality Check: Why Self-Generated AI Capabilities Don't Work
New research reveals a massive gap between AI agent marketing promises and operational reality — most self-improving agents are elaborate theater.