We are building the agent economy on a C+
The best safety grade any frontier lab earned this summer is one you would hide from your parents. The industry wiring agents into everything barely looked up.
I spent Tuesday morning with a report card. The Future of Life Institute graded nine frontier AI companies on safety this summer, and the best mark in the class was a C+. That grade went to Anthropic. OpenAI earned a C, Google DeepMind a C, Meta a D+. Three companies, xAI, DeepSeek and Mistral, failed outright.
Read the list again and remember what these companies actually are. They are not nine random startups. They supply nearly every model powering the agent economy this publication covers. Every autonomous agent resolving support tickets, moving inventory, or writing production code runs on intelligence that just graded out somewhere between mediocre and failing.
The F nobody wants to talk about
The domain scores are worse than the headlines. On existential safety, the panel's term for planning around systems that could slip meaningful human control, no company cleared a C-. Six of the nine took an F. Stuart Russell, who sat on the review panel, put it plainly: companies have backed away from earlier commitments to release new systems only with safety measures matched to their capability levels.
That last part is the piece I keep returning to. Grades can improve. A retreat from red lines is a direction, and directions compound. Two years ago most labs said publicly that certain capability thresholds would trigger a pause. The Summer 2026 index found those commitments weakened or quietly dropped across the board, right as the same labs ship models built to act on their own.
Agents raise the price of a bad grade
Here is why I think this matters more for our readers than for the general AI audience. A chatbot that fails says something wrong. An agent that fails does something wrong. It moves the money, sends the email, changes the record. We covered the BeSafe benchmark, where the best agent completed risky tasks safely about 35 percent of the time. Stack that on a C+ and you get an industry busy deploying hands and feet for systems whose makers cannot manage a B on the basics.
What almost everyone overlooks is procurement. I have read agent vendor questionnaires that ask about SOC 2, data residency and uptime, and never once ask which models sit underneath or what those models scored on any public safety evaluation. Buyers treat the model layer like electricity. It is not electricity. It varies by vendor, it changes monthly, and as of this week it has a report card anyone can read.
What I would actually do
I am not arguing that anyone should stop building. This publication exists because the agent economy is real, and mostly good news for the people who read us. I am arguing that the grade should show up in the deal. Ask your platform vendor which models run your workloads. Put model identity and substitution rights in the contract. If a vendor swaps a C model for an F model to save money, you should learn about it from a changelog, not an incident report.
The C+ itself does not scare me. Young industries grade poorly on their first report cards, and the Future of Life Institute built this index precisely so the grades would have somewhere to go. What scares me is the shrug. Nine companies were graded on whether the most consequential technology of the decade is being built carefully, the best of them cleared a bar you would hide from your parents, and the market moved on before lunch. The panel did its job. The buying side has not started doing theirs.
Opinion columns reflect the personal views of the author. Our reporting on the stories referenced here lives on the linked pages.