AI Governance - 6 min read - 12 July 2026

Every major AI lab just got graded C or below on safety. Here's how to actually use that.

An independent panel of researchers reviewed nine leading AI developers on risk assessment, governance and safety practice. Nobody scored higher than a C+. For enterprises picking an AI vendor, that unglamorous result is more actionable than any capability benchmark.

The Future of Life Institute published its Summer 2026 AI Safety Index this month, and the headline number is stark: across nine major AI developers, the highest grade awarded was a C+, given to Anthropic. OpenAI and Google DeepMind each received a C, Meta a D+, Z.ai and Alibaba Cloud a D-, and xAI, DeepSeek and Mistral all received an outright F. Time's coverage and the Institute's own report both make the same point: no company earned an A or a B, and three of the nine failed outright, one each from the US, China and Europe, which undercuts any argument that this is a story about one country's labs being reckless and everyone else's being responsible.

The grades came from an independent panel of seven researchers and governance specialists, who reviewed each company against six categories: risk assessment, current harms, safety frameworks, existential safety, governance and accountability, and information disclosure. Evidence collection ran up to 3 June 2026 and combined public material with a company survey, so the grades reflect what these organisations actually publish and disclose, not what their marketing teams say in a press release. Separately, Axios reported that several labs which had previously ruled out military applications, including some of the higher-graded companies, have since walked that commitment back and begun pursuing defence contracts, which the Institute flagged as part of a broader pattern of safety pledges softening as the commercial pressure to compete has intensified.

Why a mediocre index is more useful than a glowing one

If every major lab had scored well, the index would tell buyers almost nothing, because it wouldn't discriminate between vendors. A spread from C+ down to F, built on documented evidence rather than self-reported claims, does something more useful: it gives procurement and risk teams an independent, comparable data point that exists completely outside the vendor's own sales process. Every AI vendor pitch includes a slide about responsible AI. Very few enterprises have the in-house expertise, or frankly the standing, to independently verify what's behind that slide. An index built by outside researchers, using a consistent methodology applied to every company, is exactly the kind of third-party evidence that a normal vendor risk process is built to look for in other domains, and has mostly been missing for frontier AI.

What the categories actually tell you

The six categories map fairly cleanly onto questions a sensible AI vendor questionnaire should already be asking. Risk assessment and current harms speak to whether a vendor is honest with itself, and with customers, about what can go wrong with its models today, not just in speculative future scenarios. Safety frameworks and governance and accountability speak to whether responsible deployment is a structured, resourced function inside the company or a slide deck. Information disclosure speaks to something very practical for an enterprise buyer: how much you'll actually be told, in an incident, versus how much you'll have to infer. A vendor's grade in each of these areas is a reasonable proxy for how that vendor will behave under pressure, which is precisely the moment procurement due diligence is meant to protect you from.

How to fold this into vendor selection without overreacting to a single grade

The mistake would be to treat a C+ as disqualifying, or a C as functionally equivalent to an F. Grades like these compress a lot of nuance, and the gap between the top and bottom of the "passing" range matters as much as the letter itself. The more useful move is to use the index as a starting point for a specific conversation with each shortlisted vendor: which category did you score weakest in, what changed since the evidence was collected, and what would move that grade next time. A vendor that engages seriously with that question is telling you something different from one that dismisses the index as unrepresentative, regardless of which letter grade they were handed.

  • Add the Future of Life Institute's AI Safety Index, or an equivalent independent assessment, as a standing input to AI vendor due diligence, alongside security and compliance questionnaires.
  • Ask shortlisted vendors directly about their weakest-scoring category and what concrete steps they've taken since the evidence window closed.
  • Weight information disclosure and governance and accountability scores heavily for any vendor whose model will sit inside a regulated or customer-facing workflow.
  • Don't treat a single grade as disqualifying; treat a pattern of declining or stagnant scores across successive index editions as the real warning sign.
  • Revisit vendor scores at contract renewal, not just at initial selection, since this index is published on a roughly six-month cycle and grades move.

An index that hands out C-pluses instead of gold stars won't make for an exciting vendor slide, which is exactly why it's worth paying attention to. It's one of the few pieces of AI vendor evidence in circulation that wasn't written by the vendor. Want help building an AI vendor due diligence process that goes beyond the sales deck? Email sales@halfteck.com.

Explore more resources

Browse our full library of enterprise cloud, software, data and AI content.

View all resources