METR — short for Model Evaluation and Threat Research, pronounced “Meter” — said there was no indication that sensitive information was obtained. The attacks have not been attributed to a known group or other known threat actor. According to the company, they did not involve AI agents that violated its assessment procedures.
Theft of the API key
In March, a METR researcher without access to sensitive data was running agents on a personal EC2 instance. The instance was intentionally made available online and protected by Google authentication. The system contained an API key for the METR account on publicly available models.
The application was built using so-called “vibe coding” and contained a fail-open vulnerability. This caused authentication to be disabled without any warning, leaving the agent dashboard exposed online for several days.
METR believes the attacker discovered the system by searching for recently registered websites, including certificate transparency lists. He was likely looking for applications with references to LLMs or agents, which could contain exposed API keys from model providers.
After locating the system, he directly instructed an agent to reveal the API key. He then added his own SSH key to maintain access, and used the stolen credentials to consume API credits on public models for about three weeks.
The credits would have been worth about $600.000 if the model provider had not offered them to METR for free. The company did not disclose who the provider was.
The malicious use was not immediately detected because METR assessments and experiments typically consume a large number of tokens. Additionally, the API keys had no spending limits. Following the incident, the organization amended its policies on the use of credentials and data on non-METR infrastructure or devices, improved monitoring, and added spending alerts where possible.
The May campaign
The second incident was recorded in May 2026 and was described as a sustained external attack campaign. METR believes that it was likely carried out by a financially motivated actor who was attempting to gain illicit access to advanced AI models.
The attackers conducted systematic testing on METR's public infrastructure and heavily utilized agents to automate vulnerability scanning. Among other things, they attempted credential stuffing on authentication providers, attempted to obtain OAuth tokens, scanned newly deployed services, and attempted to phish staff members.
At the same time, METR discovered that it had accidentally exposed a read-only SQL query mechanism that existed in the public transcript viewer. Queries were normally limited to public data, but a bug in this component could have allowed access to unpublished assessment data.
At the same time, the database accidentally contained sensitive model data, even though it was supposed to only store data from non-sensitive models. The issue was discovered when an independent security researcher reported it to METR. The organization then disabled the API.
The attackers briefly looked at the endpoint as part of the broader campaign, but METR says there is no indication they discovered the vulnerability or gained access to non-public data.
Your comments will not be published if: