Free Consultation
Observatory

Show your work

We argue everywhere else on this site that a governance claim should be checkable. This page is that standard turned on ourselves — what our agents changed without asking a human, what our own integrity agent says is wrong with us right now, every prediction this blog has committed to, and what we got wrong. Nothing here is curated for how it looks.

3
open findings against ourselves
none critical or high
no prediction settled yet
10 on the ledger · 1 past due
100%
of incident candidates rejected
126 judged, 0 kept
0
changes agents published with no human
0 in the last 7 days
2
official links we publish that are broken
out of 38 checked

The integrity agent last ran 2h ago. It runs every six hours against the live database and against this repository's own source files, and it has no way to mark itself healthy — it can only fail to find something.

What our own agent says is wrong with us

Open findings only, exactly as the agent filed them. Findings we have acknowledged or already fixed are not shown, because listing them here would pad this section with our own good news. The evidence payloads and the fix buttons stay in the admin dashboard; the statement of what is wrong is public.

medium Data going stale open since 2026-08-21

1 regulation source has changed since we last read them

The official text at the URL we publish has moved materially for: Algorithm, Deep Synthesis & Generative AI Measures. Our status prose was written from the previous version, so it may now be wrong on a public page. Re-read the source and update regulations.js — the agent deliberately does not rewrite the catalogue itself, because a running server editing its own source would not survive a redeploy.

low A backlog building up open since 2026-08-21

Prediction due 2026-06 still has no verdict (2 months on)

We published this and put it on the public ledger at /proof: "By June 2026, at least one major AI incident in an under-reported category (e.g., environmental harm or overreliance) will be revealed through a non-traditional source (e.g., academic paper, whistleblower, or citizen report) that the trade ". Its date has passed. What would have proved it wrong: No such incident emerges, and the trade press continues to be the primary source for all consequential AI incidents.. Settle it as came true, wrong, or too unclear to call — "too unclear" is a legitimate verdict and an admission the prediction was not written sharply enough, which is more useful published than hidden. 1 prediction(s) are past due in total.

low Data going stale open since 2026-08-21

5 regulation sources refuse automated checks

These hosts answer 401/403/429 to any non-browser request, so the source watch can never verify them: ca_qc_law25, ca_standards, iso_23894, iso_42001, us_co_sb21_169. The links are not broken — they work in a browser — but these entries stay human-verified only. Nothing to fix unless a host publishes a machine-readable alternative.

The prediction ledger

Every original-thesis post this blog publishes has to commit to a dated prediction and say what would prove it wrong, or it does not get published. All of them are here, settled or not.

We have no hit rate yet. Nothing on this ledger has been settled, so there is no honest number to report — and a hit rate computed only over the ones we chose to settle would not be one either. 1 prediction is past due with no verdict, the oldest from 2026-06. Our own integrity agent raises a finding for each one until it is settled.

Past due, no verdict said it would be true by 2026-06

By June 2026, at least one major AI incident in an under-reported category (e.g., environmental harm or overreliance) will be revealed through a non-traditional source (e.g., academic paper, whistleblower, or citizen report) that the trade press had missed entirely.

The argument it came from: The concentration of AI incident reporting in just five legal and security trade outlets means that the AI risk landscape is being shaped by the legal profession's and security vendors' interests, not by actual system failures, and this is systematically hiding the most consequential risks like overreliance and environmental harm.

What would prove it wrong: No such incident emerges, and the trade press continues to be the primary source for all consequential AI incidents.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, at least one major EU-based AI developer will publicly announce a reduction in generative AI features for EU users, citing the NCII/CSAM ban's compliance burden, and will relocate development to a non-EU jurisdiction.

The argument it came from: The EU AI Act's NCII/CSAM ban (effective 2026-12-02) will create a 'false positive crisis' for legitimate AI developers, because the ban's strict liability for generating prohibited content will force developers to over-filter outputs, and the incident data shows that 'genai' stories have dropped by 8 in the last 45 days—evidence that developers are already self-censoring, which will suppress legitimate use cases and drive innovation offshore.

What would prove it wrong: No such announcement occurs, and EU-based generative AI usage and development continues to grow at the same rate as before the ban.

Read the post that made this call →
Open said it would be true by 2027-02

By 2027-02, the average salary for an AI model-risk validator in North America will have risen by at least 25%, and at least 20% of OSFI-regulated institutions will publicly warn of compliance delays.

The argument it came from: The OSFI Guideline E-23, effective 2027-05, will force Canadian financial institutions to treat AI models as regulated 'models' under existing model-risk management, and this will create a talent bottleneck that drives up compliance costs by 40% for early adopters, a cost that is not being budgeted for.

What would prove it wrong: If salaries rise less than 10% or no institution warns of delays, the talent bottleneck is overstated.

Read the post that made this call →
Open said it would be true by 2027-01

By 2027-01, at least one major AI governance framework will add 'memory manipulation' or 'agent tool invocation' as a named risk category, directly borrowing from security research.

The argument it came from: The security community is 18 months ahead of the governance community on AI risk, and the gap is not a lag but a structural misalignment: security writes about actively exploited vulnerabilities and supply chain attacks, while governance writes about privacy and fairness, meaning the risk teams are defending against yesterday's threats while the attackers have already moved to agentic and memory-based attacks.

What would prove it wrong: If governance frameworks continue to ignore these attack vectors, the structural gap claim is overstated.

Read the post that made this call →
Open said it would be true by 2027-04

By 2027-04-01, legal commentary and CPPA guidance will reveal that companies with large language models are claiming they are exempt from ADMT opt-out requirements by asserting 'human-in-the-loop review,' while companies with simpler automated scoring systems are being forced to comply—and the ratio of ADMT exemption claims for neural models vs. rule-based systems will be at least 3:1 in published compliance guidance.

The argument it came from: The 2027 California ADMT opt-out and pre-use notice requirements will create a perverse incentive for companies to deploy less transparent, less explainable AI models—because a model that cannot articulate why it made a decision is easier to claim is 'not subject to ADMT' than a rule-based model whose logic is auditable and therefore demonstrably automated.

What would prove it wrong: If the CPPA's final ADMT regulations explicitly state that any model that generates a decision affecting a consumer is subject to opt-out regardless of human review, and companies with black-box models are found to be complying at the same rate as those with transparent systems, this claim is wrong.

Read the post that made this call →
Open said it would be true by 2027-06

By 2027-06, at least one major AI governance framework will explicitly incorporate MITRE ATLAS techniques into its risk taxonomy, and public incident reporting will start citing these techniques.

The argument it came from: The attack techniques with documented real-world cases (e.g., LLM Prompt Crafting, 22 cases) are absent from the news window because the governance world is focused on privacy and fraud, not on the adversarial techniques that security teams are actively tracking—and this gap means governance frameworks are being built on the wrong threat model.

What would prove it wrong: If no major framework or regulator references ATLAS techniques by then, the misalignment is not being corrected.

Read the post that made this call →
Open said it would be true by 2028-08

By 2028-08, when the EU AI Act Annex I embedded high-risk obligations apply, at least two major industrial AI deployments will face sudden regulatory halts due to unmeasured robustness failures that never appeared in prior incident feeds.

The argument it came from: The absence of reporting on environmental harm and system robustness is a structural failure of incident databases, not an indicator of zero risk.

What would prove it wrong: If all regulatory halts or enforcement actions under EU AI Act Annex I prior to August 2028 stem exclusively from risks already tracked in current top-five incident categories.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, public governance incident reports will continue to decline even as total internal AI incidents (as measured by corporate disclosures in securities filings) rise.

The argument it came from: The steep decline in governance-related AI incident reporting is not a sign of improvement but a leading indicator that organizations are shifting incident disclosure from public channels to private legal and compliance workflows ahead of enforceable deadlines.

What would prove it wrong: If public governance incident reports increase while compliance reports also rise, the shift hypothesis is wrong.

Read the post that made this call →
Open said it would be true by 2027-03

By early 2027, at least two major AI incidents will be revealed to have been reported as 'compliance issues' in 2026, with the underlying vulnerability still unpatched.

The argument it came from: The collapse in reported AI incident volume (from 450 to 344 stories) is not a decline in risk but a migration of incidents into the 'compliance' category, indicating that organizations are re-labeling failures to fit regulatory checklists rather than fixing underlying system flaws.

What would prove it wrong: If a retrospective audit of 2026 incident reports shows no systematic re-labeling from governance to compliance categories.

Read the post that made this call →
Open said it would be true by 2028-01

Publicly disclosed internal AI algorithm failures by Fortune 500 companies operating in California will decline by at least 25% in the twelve months following the January 2027 CPPA effective date.

The argument it came from: The impending enforcement of the California CPPA ADMT rules in early 2027 will cause enterprises to aggressively suppress internal AI incident logging to avoid creating discoverable compliance evidence.

What would prove it wrong: If public disclosures of internal AI algorithm failures by Fortune 500 companies increase or remain flat through January 2028.

Read the post that made this call →

What the agents changed with no human

When the official text behind a regulation moves, an agent reads it and drafts the new status. It only publishes if three independent readings agree on the confidence, on the status category, and on whether the change is material at all. Any dissent and nothing is published. Every change it did publish is below, with the before and after.

No agent has published a change on its own yet. This stays empty until three independent readings agree on one — which is the bar, and most candidates never clear it.

What we threw away

The promoter has judged 126 candidates from the live news feed and kept 0. It last judged something 2h ago, with 728 waiting.

100% of everything it looked at was rejected. That is the intended outcome — most AI news is commentary, not an incident — and the reasons are below.

  • 90 not actually an incident
  • 36 too thin to write up

Links we publish that are broken

Every regulation page links its official text. A watcher re-fetches all of them on a six-hour cycle, so a dead one is our broken promise and gets listed here rather than quietly sitting on the page. A further 5 hosts refuse automated checks entirely; those are verified by hand and are not counted as broken.

dead link last checked 8h ago

Illinois HB 3773 — AI in Employment (IHRA amendment)

https://www.ilga.gov/legislation/BillStatus.asp?DocNum=3773&GAID=17&DocTypeID=HB&SessionID=112

Remote server returned HTTP 404.

The page that publishes this link →
dead link last checked 2h ago

Biometric Information Privacy Act (BIPA)

https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004

Remote server returned HTTP 404.

The page that publishes this link →

How to check any of this

  • Every regulation page carries the date its own official source was last read, and warns you when it has moved since we wrote the summary — start at the regulation reference.
  • Every curated incident says what it was, what failed, and the control that would have prevented it — the incident radar.
  • Every prediction above links the post that made it. The posts are dated and were not edited afterwards.
  • Every dated obligation we publish has its own page and a calendar feed you can subscribe to and check against the official source yourself.
  • Methodology & Scoring gives the actual formula behind every score and classification on this site, so you can recompute one rather than trust it.