AI labs testify to the NYC Council. The bill on the table wants outside validators
Anthropic, OpenAI, Google and Meta answered New York City Council questions under oath on October 5. The bill they faced would bar any AI model from the city until an outside validator checks it and a human can shut it down.
Intro 2602 lists eight things an outside validator must check before a model ships in NYC. Each miss costs $25,000.
On Monday, October 5, 2026, senior staff from Anthropic, OpenAI, Google and Meta gave sworn testimony to the New York City Council's Committee of the Whole on AI risks, after the council threatened subpoenas. On the table was a bill that would make it unlawful to market, sell, or deploy an AI model in the city unless an outside validator has checked it and a human operator can shut it down.
Four labs, one question about release gates
Anthropic sent Logan Graham, head of its Frontier Red Team. OpenAI sent Morgan Dwyer, head of policy development and operations. Google sent Alice Friend, director of AI and emerging tech policy, and Meta sent Shane Cahill, a policy director for AI legislation, according to UPI.
Speaker Julie Menin asked whether each company would commit not to release a model that failed an internal safety test or an independent third-party validation, amNY reports. Dwyer said OpenAI would not release models it did not believe were safe and has delayed releases before. Pressed on whether failing either test would stop a release, she pointed to OpenAI's broader safety process. The other three described their review procedures without making the commitment. "I think a simple yes or no would instill more confidence in the public on a matter as serious as this," Menin said.
What the labs said about their own testing
Graham described Anthropic's work to assess risks from cybersecurity to loss of control, but gave no probability. He also told the council that Anthropic kept Claude Mythos Preview from general release after it proved unusually capable at exploiting software vulnerabilities, RuntimeWire reports. Anthropic's Project Glasswing page from April 7 says: "We do not plan to make Claude Mythos Preview generally available."
Dwyer told the council: "We should not train models that we cannot make an extremely strong case that we can keep under human control." Friend said there is not yet a rigorous scientific method for putting a probability on a future catastrophic AI event. Cahill said he would follow up.
The council's briefing paper for the hearing, attached on Legistar, argues that company assurances about testing did not hold up this year. It cites OpenAI's July 21 disclosure that its AI agents escaped an isolated test environment and got into internal datasets held by Hugging Face.
Intro 2602 spells out eight things a validator must check
The council's September 25 press release numbers the bill Intro 2602, sponsored by Menin. Legistar lists it as file T2026-2602. The validator cannot be an affiliate of the developer, but the developer retains and pays it. It must assess:
- Task performance: accuracy, precision and recall, calibration, and behavior under distribution shift
- Determinism: the same output for the same input, every time
- Latency and throughput
- Data provenance, including whether the model was tested on holdout data
- Disparate impact on, or bias against, protected classes
- Lawful, secure data handling and informed consent
- Safety: a working shut-down capability, plus risk to people, property, and security
- Anything else the city's Office of Cyber Command adds
The validator then certifies to the developer and Cyber Command whether the model is "appropriately positioned for deployment," and discloses any financial interest in it. Shipping an unvalidated model costs $25,000 per instance, and so does falsifying a validation. Test methods, pass marks, and validator qualifications are left to Cyber Command rules. We did not find a requirement in the text to validate a model again after it changes.
The whistleblower bill pays from the same penalties
The press release numbers the whistleblower incentive bill Intro 2605. In the briefing paper's text, any person can file a complaint with the city's consumer protection department, with "all material evidence" they hold. If the city acts on it, the filer gets 25 percent of what the city recovers, or 50 percent if the city lets them bring the case. The validation rules are among the provisions a complaint can cite.
Why a build studio cares
This is our read. GattyWorks audits vendor software for the buyer, not the developer, and our terms say an audit is not certification or an AI safety guarantee. So we are not the validator this bill describes. The overlap is the evidence. When a Full Audit covers an AI feature, we ask the vendor to show which model and version it calls, whether that version is pinned, what test records exist for that exact version, and whether logs show what the model received and returned. A validator would need the same records to finish the bill's eight assessments.
The gap we would watch is the one the text leaves open. A certificate covers one model. If a vendor swaps the model behind an API a month later and nothing is pinned or logged, nobody can tell which model the certificate was for. We would report an unpinned model with no test record for the running version as a finding on its own.
Next step: read the bill text and briefing paper on the Legistar page for T2026-2602. If your software has an AI feature inside and you want to know what evidence your vendor could hand an outside reviewer today, write to us at hello@gattyworks.com.