Browse all tools and resources →

Read me Page help ↗
Observatory

Show your work

We argue everywhere else on this site that a governance claim should be checkable. This page is that standard turned on ourselves — what our agents changed without asking a human, what our own integrity agent says is wrong with us right now, every prediction this blog has committed to, and what we got wrong. Nothing here is curated for how it looks.

4
open findings against ourselves
1 of them critical or high
—
no prediction settled yet
45 on the ledger · 1 past due
100%
of incident candidates rejected
1379 judged, 0 kept
0
changes agents published with no human
0 in the last 7 days
2
official links we publish that are broken
out of 38 checked

The integrity agent last ran 3h ago. It runs every six hours against the live database and against this repository's own source files, and it has no way to mark itself healthy — it can only fail to find something.

What our own agent says is wrong with us

Open findings only, exactly as the agent filed them. Findings we have acknowledged or already fixed are not shown, because listing them here would pad this section with our own good news. The evidence payloads and the fix buttons stay in the admin dashboard; the statement of what is wrong is public.

high Data going stale open since 2026-08-21

5 regulation sources have changed since we last read them

The official text at the URL we publish has moved materially for: Algorithm, Deep Synthesis & Generative AI Measures; MAS FEAT Principles + AI Model Risk Management; Fair Credit Reporting Act (FCRA); Digital Operational Resilience Act (DORA, Reg. 2022/2554); Title VII / ADA applied to AI hiring tools (EEOC). Our status prose was written from the previous version, so it may now be wrong on a public page. Re-read the source and update regulations.js — the agent deliberately does not rewrite the catalogue itself, because a running server editing its own source would not survive a redeploy.

medium Data going stale open since 2026-09-06

2 regulation source check(s) could not complete

Timeouts, connection failures and server errors do not prove that a link is dead. Source watch retries automatically with backoff. Check the source and retry time below; investigate persistent failures. Acknowledge keeps the same condition in Reviewed; a new affected source or escalation reopens it.

low Data going stale open since 2026-08-21

4 regulation source(s) need manual access checks

These requests were refused or rate-limited (HTTP 401, 403 or 429), or returned an access-challenge page. This does not establish whether the link works in a browser. Open each official source to review it. Automatic retries continue with backoff. Acknowledge moves this unchanged condition to Reviewed; it does not verify the regulation.

low A backlog building up open since 2026-08-16

22 regulatory candidate(s) unreviewed for over 14 days

These sit in the Regulatory Watch queue waiting on a human. Open the admin dashboard’s Regulatory review & platform operations → Regulatory Watch to assess the candidates. Age alone does not mean an item is irrelevant. Use the bulk dismiss action only after review; it leaves anything already promoted to the Horizon unchanged.

The prediction ledger

Every original-thesis post this blog publishes has to commit to a dated prediction and say what would prove it wrong, or it does not get published. All of them are here, settled or not.

We have no hit rate yet. Nothing on this ledger has been settled, so there is no honest number to report — and a hit rate computed only over the ones we chose to settle would not be one either. 1 prediction is past due with no verdict, the oldest from 2026-06. Our own integrity agent raises a finding for each one until it is settled.

Past due, no verdict said it would be true by 2026-06

By June 2026, at least one major AI incident in an under-reported category (e.g., environmental harm or overreliance) will be revealed through a non-traditional source (e.g., academic paper, whistleblower, or citizen report) that the trade press had missed entirely.

The argument it came from: The concentration of AI incident reporting in just five legal and security trade outlets means that the AI risk landscape is being shaped by the legal profession's and security vendors' interests, not by actual system failures, and this is systematically hiding the most consequential risks like overreliance and environmental harm.

What would prove it wrong: No such incident emerges, and the trade press continues to be the primary source for all consequential AI incidents.

Read the post that made this call →
Open said it would be true by 2028-03

At least one EU or US data protection authority will open a formal inquiry, or issue guidance, that treats model-inferred sensitive attributes (not merely collected data) as the basis of the violation — and at least one enforcement decision will cite inference rather than collection as the harm.

The argument it came from: The privacy incident surge is not a privacy problem — it is the first measurable evidence that AI systems are being deployed as inference engines over data that was lawfully collected and is now being unlawfully derived from, and the correct control is a derivation register, not a consent refresh.

What would prove it wrong: If by 2028-03 every privacy enforcement action in the EU/UK/US that touches AI cites collection, retention, or transfer as the violation, with no decision or guidance treating inference-from-lawful-data as the harm, the claim fails. Check EDPB decisions and guidance, ICO enforcement notices, and CPPA enforcement actions.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, at least one major AI risk platform, insurer, or standards body will publish an AI incident index explicitly framed as a litigation or enforcement activity index rather than a harm-frequency index, and the Riskwell feed will show a category where story count and documented-case count move in opposite directions for two consecutive 45-day windows.

The argument it came from: The 7% top-five-outlet concentration is not a coverage problem to be fixed by better journalism — it is a structural feature that makes the AI incident record useless for trend detection, and the correct governance response is to stop treating incident counts as a risk signal and start treating them as a market signal about which risks are currently litigable.

What would prove it wrong: If by 2027-03 no such litigation-framed index exists and no category shows a sustained inverse relationship between story count and documented-case count, the reinterpretation is wrong.

Read the post that made this call →
Open said it would be true by 2027-08

At least two major open-source foundation models released after December 2026 will experience documented enterprise boycotts due to excessive false-positive refusals triggered by EU compliance guardrails.

The argument it came from: The December 2026 EU AI Act ban on NCII/CSAM and marking grace period expiration will force open-source model maintainers to implement catastrophic over-filtering that degrades benign enterprise semantic search capabilities.

What would prove it wrong: Enterprise adoption rates for open-source foundation models in the EU will grow by more than 25% year-over-year throughout 2027 without reported performance regression from guardrails.

Read the post that made this call →
Open said it would be true by 2027-06

At least three major enterprise AI deployments in the EU will face formal non-compliance notices under the December 2026 AI Act CSAM rules due to prompt-crafting exploits rather than intentional data hosting.

The argument it came from: The December 2026 EU AI Act CSAM enforcement deadline will trigger a wave of false-positive governance penalties because organizations are treating illegal content filtering as a model accuracy problem rather than an agentic prompt-crafting vulnerability.

What would prove it wrong: If European regulators issue zero enforcement actions or warning letters regarding automated content moderation bypasses under the Digital Omnibus rules by June 2027.

Read the post that made this call →
Open said it would be true by 2027-08

By 2027-08, at least one major agentic AI vendor will disclose a security incident where an API key compromise led to cascading access across multiple downstream systems, and this incident will be the first to be reported in both the security and governance news feeds simultaneously, marking the convergence of the two communities.

The argument it came from: The security community's top concern — 'Your AI's API Key Is a Master Key' and 'Supply Chain Attack' (6 articles) — points to a risk that governance frameworks do not address: AI systems inherit the access rights of the tools they invoke, so a single compromised API key in an agentic workflow can trigger cascading failures across multiple systems, and this is a governance gap that will be exploited before the 2028 EU AI Act embedded high-risk obligations take effect.

What would prove it wrong: If no such cascading API-key incident is disclosed by mid-2027, and the governance and security feeds remain disjoint on this topic, then either the risk is lower than the security community believes or the governance community is even more disconnected than argued.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, at least one major EU-based AI developer will publicly announce a reduction in generative AI features for EU users, citing the NCII/CSAM ban's compliance burden, and will relocate development to a non-EU jurisdiction.

The argument it came from: The EU AI Act's NCII/CSAM ban (effective 2026-12-02) will create a 'false positive crisis' for legitimate AI developers, because the ban's strict liability for generating prohibited content will force developers to over-filter outputs, and the incident data shows that 'genai' stories have dropped by 8 in the last 45 days—evidence that developers are already self-censoring, which will suppress legitimate use cases and drive innovation offshore.

What would prove it wrong: No such announcement occurs, and EU-based generative AI usage and development continues to grow at the same rate as before the ban.

Read the post that made this call →
Open said it would be true by 2027-12

By December 2027, internal audit reviews of companies complying with the January 2027 California CPPA and Colorado ADMT deadlines will reveal that over 40% of their logged 'privacy incidents' were actually undocumented robustness or overreliance failures.

The argument it came from: The 42% surge in privacy incidents acts as a regulatory camouflage that actively conceals zero-reported overreliance and lack of robustness risks.

What would prove it wrong: If public enforcement actions or independent audits of CPPA/ADMT compliance by December 2027 show fewer than 10% of filed privacy-labeled incidents involved underlying robustness or overreliance root causes.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, internal whistleblowers or regulatory discovery will reveal that at least two major enterprises systematically reclassified safety incidents as privacy leaks.

The argument it came from: The steep drop in safety reporting is an artifact of corporate compliance shifting incidents into unmonitored privacy silos.

What would prove it wrong: Regulatory filings and transparency reports through March 2027 show zero instances of reclassified safety-to-privacy incidents.

Read the post that made this call →
Open said it would be true by 2027-04

Over 40% of Fortune 500 enterprise public AI incident disclosures filed in Q1 2027 will classify adversarial model extractions and prompt injections strictly under privacy compliance headings.

The argument it came from: The impending enforcement of California's ADMT and the Colorado ADMT Act on January 1, 2027 will trigger a surge in corporate privacy re-labeling, where organizations formally reclassify security vulnerabilities and prompt injection exploits as internal privacy breaches to satisfy compliance deadlines without fixing model robustness.

What would prove it wrong: If public disclosures filed with the California CPPA in Q1 2027 explicitly categorize more than 50% of AI-related incidents under security vulnerability or robustness codes rather than privacy.

Read the post that made this call →
Open said it would be true by 2028-06

By 2028-06, a named EU regulator or industry survey will publicly report that over 30% of AI incidents classified as 'minor' by firms were later reclassified as 'serious' upon review.

The argument it came from: The EU AI Act's 2027-12-02 serious-incident reporting obligation will be gamed by firms classifying incidents as 'minor' to avoid reporting, mirroring the 50% drop in critical incidents already visible in the data.

What would prove it wrong: If no EU regulator or credible industry survey publishes such a reclassification rate by 2028-06, the claim is wrong.

Read the post that made this call →
Open said it would be true by 2027-12

Enterprise legal spend on AI compliance documentation will grow by over 50% while internal technical audits of AI models decrease.

The argument it came from: The divergence between compliance stories rising and governance stories plunging proves that enterprises are treating regulatory deadlines as paperwork exercises rather than security upgrades.

What would prove it wrong: If enterprise budget allocations show technical audit spending growing faster than legal documentation spending in annual risk surveys through 2027

Read the post that made this call →
Open said it would be true by 2027-11

By November 2027, at least three major Canadian financial institutions will face regulatory censure for validating generative AI systems using static MRM pipelines that missed robustness failures.

The argument it came from: OSFI Guideline E-23 taking effect in May 2027 will force financial institutions to apply traditional banking Model Risk Management frameworks to LLMs, breaking entirely because models are non-deterministic.

What would prove it wrong: If OSFI issues an explicit exemption or relaxed validation standard for generative AI models prior to November 2027.

Read the post that made this call →
Open said it would be true by 2027-12

By 2027-12-02, at least one major enterprise enforcement action under EU AI Act Annex III will target a multi-agent orchestration failure where no single vendor or deployer can be assigned sole accountability.

The argument it came from: Multi-agent systems will become the primary vector for untraceable corporate liability because current governance frameworks treat agents as single-instance tools rather than distributed supply chains.

What would prove it wrong: If all EU AI Act Annex III enforcement actions prior to this date name single-vendor, monolithic AI systems without multi-agent handoffs.

Read the post that made this call →
Open said it would be true by 2027-02

By 2027-02, the average salary for an AI model-risk validator in North America will have risen by at least 25%, and at least 20% of OSFI-regulated institutions will publicly warn of compliance delays.

The argument it came from: The OSFI Guideline E-23, effective 2027-05, will force Canadian financial institutions to treat AI models as regulated 'models' under existing model-risk management, and this will create a talent bottleneck that drives up compliance costs by 40% for early adopters, a cost that is not being budgeted for.

What would prove it wrong: If salaries rise less than 10% or no institution warns of delays, the talent bottleneck is overstated.

Read the post that made this call →
Open said it would be true by 2028-12

By December 2028, over 30% of critical operational stoppages involving enterprise AI agents will be attributed to systemic robustness degradation rather than security breaches or prompt attacks.

The argument it came from: The complete absence of 'Lack of capability or robustness' stories in governance databases proves that enterprise risk registers are failing to measure operational degradation under load.

What would prove it wrong: If post-incident analyses published by major risk tracking services attribute less than 15% of operational failures to robustness degradation by end of 2028.

Read the post that made this call →
Open said it would be true by 2028-06

At least one major enterprise deployment in a regulated US state or EU jurisdiction will be halted or fined due to non-compliance with regional compute-energy thresholds.

The argument it came from: The complete absence of environmental harm reporting in mainstream governance outlets will leave organizations entirely unprepared for imminent grid-capacity liability under upcoming AI regulations.

What would prove it wrong: Regulatory filings showing zero enforcement actions or operational restrictions related to AI energy consumption or grid capacity limits.

Read the post that made this call →
Open said it would be true by 2027-09

Between 2026-12 and 2027-09, at least one civil complaint or regulatory action in the EU or US will cite a synthetic-media provenance mark or watermark as the basis for attributing an output to a named provider, and at least one major model provider will ship a mark that is not robust to common transformations.

The argument it came from: The EU AI Act's 2026-12-02 end of the marking grace period will convert every unmarked synthetic output into a strict-liability artifact, and the second-order effect is that provenance marking will become an admission mechanism: the same watermark that satisfies the Act will let plaintiffs and regulators attribute outputs to a specific model version, which will push firms toward weaker or removable marking and make the compliance control self-defeating.

What would prove it wrong: If by 2027-09 no public complaint or enforcement action cites a provenance mark as attribution evidence, the liability-conversion thesis is wrong and marking is purely a compliance cost.

Read the post that made this call →
Open said it would be true by 2027-12

Over 50% of major enterprise AI outages cited in regulatory filings during the second half of 2027 will be attributed to multi-agent loops rather than single-model hallucination or lack of robustness.

The argument it came from: The divergence between zero public robustness stories and 109 multi-agent incident reports proves that organizations are auditing model accuracy while operating blind to cascading agent coordination failures.

What would prove it wrong: If public post-mortems and regulatory disclosures through December 2027 show that single-model robustness failures outnumber multi-agent coordination failures by a ratio of 2:1 or higher.

Read the post that made this call →
Open said it would be true by 2027-12

Regulatory bodies will issue guidance specifically addressing the reclassification of security failures as privacy incidents by December 2027.

The argument it came from: The decline in reported security incidents is not due to improved security but because successful attacks are increasingly being reclassified as privacy incidents for regulatory compliance reasons.

What would prove it wrong: If no such guidance is issued by any major regulatory body by December 2027.

Read the post that made this call →
Open said it would be true by 2028-08

By 2028-08, the first major regulatory enforcement action under EU AI Act Annex III high-risk rules will target a robustness or overreliance failure rather than a data privacy leak.

The argument it came from: The absolute silence around model robustness and overreliance in incident feeds is a false-security artifact created by the fact that top-tier outlets prioritize privacy leaks over silent capability failures.

What would prove it wrong: If all EU AI Act Annex III enforcement actions prior to August 2028 are exclusively privacy or data governance violations.

Read the post that made this call →
Open said it would be true by 2027-01

By 2027-01, at least one major AI governance framework will add 'memory manipulation' or 'agent tool invocation' as a named risk category, directly borrowing from security research.

The argument it came from: The security community is 18 months ahead of the governance community on AI risk, and the gap is not a lag but a structural misalignment: security writes about actively exploited vulnerabilities and supply chain attacks, while governance writes about privacy and fairness, meaning the risk teams are defending against yesterday's threats while the attackers have already moved to agentic and memory-based attacks.

What would prove it wrong: If governance frameworks continue to ignore these attack vectors, the structural gap claim is overstated.

Read the post that made this call →
Open said it would be true by 2028-06

By 2028-06, at least one major incident involving 'Overreliance and unsafe use' (5.1) or 'Lack of capability or robustness' (7.3) will occur and be widely reported in mainstream outlets, forcing a re-evaluation of current AI governance frameworks.

The argument it came from: The near-total absence of reporting on Overreliance and unsafe use (5.1) and Lack of capability or robustness (7.3) risks, despite documented attack techniques targeting these weaknesses, indicates a critical blind spot in both AI governance frameworks and executive risk perception.

What would prove it wrong: If by 2028-06, no major incident related to risk classes 5.1 or 7.3 appears in the top five reporting outlets (Biometric Update, Law360, Yahoo Finance, Help Net Security, Reuters) according to the LIVE AI INCIDENT DATABASE.

Read the post that made this call →
Open said it would be true by 2028-09

At least one of the four near-zero categories (overreliance, environmental harm, capability/robustness, competitive dynamics) will be the subject of a named regulator's or standards body's formal consultation or guidance document before 2028-09, while remaining under 10 stories in our 180-day feed at that date.

The argument it came from: The five top outlets carrying only 8% of incident coverage means no outlet has enough volume to sustain a beat, so the AI incident record is now structurally incapable of detecting slow-onset harms — the exact class (overreliance, environmental harm, capability/robustness gaps, competitive dynamics) that shows 0-2 stories in 180 days — and any governance programme that treats the public record as its horizon scan is blind to precisely those risks.

What would prove it wrong: Check the consultation and guidance pages of the EU AI Office, NIST, and OSFI before 2028-09, and pull the 180-day category counts at that date. If none of the four categories has produced a formal document, or if the feed count for the covered category has risen above 10, the claim fails.

Read the post that made this call →
Open said it would be true by 2027-09

By Q3 2027, more than 50% of major AI security breaches reported in trade press will originate from compromised data science notebooks or unmanaged API keys rather than prompt injection or algorithmic bias.

The argument it came from: Security teams are tracking active exploitation of API keys and data science notebooks while governance frameworks are writing policies about fairness and bias.

What would prove it wrong: If prompt injection and algorithmic fairness issues account for the majority of reported enterprise AI breaches in public databases by Q3 2027.

Read the post that made this call →
Open said it would be true by 2027-04

By 2027-04-01, legal commentary and CPPA guidance will reveal that companies with large language models are claiming they are exempt from ADMT opt-out requirements by asserting 'human-in-the-loop review,' while companies with simpler automated scoring systems are being forced to comply—and the ratio of ADMT exemption claims for neural models vs. rule-based systems will be at least 3:1 in published compliance guidance.

The argument it came from: The 2027 California ADMT opt-out and pre-use notice requirements will create a perverse incentive for companies to deploy less transparent, less explainable AI models—because a model that cannot articulate why it made a decision is easier to claim is 'not subject to ADMT' than a rule-based model whose logic is auditable and therefore demonstrably automated.

What would prove it wrong: If the CPPA's final ADMT regulations explicitly state that any model that generates a decision affecting a consumer is subject to opt-out regardless of human review, and companies with black-box models are found to be complying at the same rate as those with transparent systems, this claim is wrong.

Read the post that made this call →
Open said it would be true by 2027-06

By 2027-06, at least one major AI governance framework will explicitly incorporate MITRE ATLAS techniques into its risk taxonomy, and public incident reporting will start citing these techniques.

The argument it came from: The attack techniques with documented real-world cases (e.g., LLM Prompt Crafting, 22 cases) are absent from the news window because the governance world is focused on privacy and fraud, not on the adversarial techniques that security teams are actively tracking—and this gap means governance frameworks are being built on the wrong threat model.

What would prove it wrong: If no major framework or regulator references ATLAS techniques by then, the misalignment is not being corrected.

Read the post that made this call →
Open said it would be true by 2027-09

The first major incident resulting from overreliance on generative AI will occur before organizations begin regularly reporting on this risk class.

The argument it came from: The underreporting of overreliance and unsafe use (0 stories in 180 days) combined with the rise in generative AI incidents (29 stories in last 45 days) suggests organizations are deploying powerful generative AI without proper safety protocols.

What would prove it wrong: If organizations begin regularly reporting on overreliance and unsafe use risks before a major incident occurs, this claim is false.

Read the post that made this call →
Open said it would be true by 2027-09

By 2027-09, at least one non-Canadian financial regulator or major standards body will publish AI model validation guidance whose structure (independent validation, documented assumptions, ongoing monitoring) tracks OSFI E-23, and our feed will carry more stories referencing OSFI E-23 than in the current window.

The argument it came from: The May 2027 OSFI E-23 model risk management guideline is the most consequential AI governance deadline on the calendar and the least discussed, because it imports banking-style model validation — independent challenge, documented assumptions, ongoing monitoring — into AI/ML at federally regulated financial institutions, and it will arrive before the EU AI Act's high-risk obligations, setting the de facto global standard for AI model validation outside the EU.

What would prove it wrong: If by 2027-09 our feed shows no stories citing OSFI E-23 as a model for other jurisdictions and no non-Canadian regulator has published comparable validation guidance, the standard-setting thesis is wrong.

Read the post that made this call →
Open said it would be true by 2027-08

Fairness-category story counts in our feed in the 45 days after 2027-04-01 exceed the 45 days before 2027-01-01 by at least 40%, with no corresponding rise in AI deployment volume attributable in the same window.

The argument it came from: Privacy stories rising 38 while fairness stories fall 22 is not a shift in what AI systems do — it is a shift in which harm has a statutory hook. Privacy has GDPR, CCPA, and an enforcement apparatus; fairness has proposed bills. Incident reporting follows enforceability, so the 'fairness decline' is a measurement artifact that will reverse the moment Colorado and California ADMT obligations bite in 2027.

What would prove it wrong: If fairness story counts in the 45 days following 2027-04-01 are at or below the level of the 45 days before 2027-01-01, the enforceability-drives-reporting claim is wrong.

Read the post that made this call →
Open said it would be true by 2027-12

The share of stories in our feed classified critical will remain below 12% of total volume through the EU AI Act Art. 73 effective date, and the privacy category will stay at or above 90 stories per 45-day window

The argument it came from: The collapse in reported critical AI incidents (83 to 48) while total volume rose is not a safety improvement but the signature of a reporting channel change: critical events are being routed into privacy breach notification workflows rather than AI incident processes, because privacy has a legal clock and AI incidents do not.

What would prove it wrong: If in any 45-day window before 2027-12-02 critical stories exceed 15% of total volume while privacy stays below 70, the intake-reclassification mechanism is wrong

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, at least one major breach involving AI agent tool invocation (where an attacker, not a user, triggered an agent to exfiltrate data or transfer funds) will be disclosed in security press, and it will be initially misclassified as a 'multi-agent failure' or 'user error' in at least one governance-oriented publication.

The argument it came from: The MITRE ATLAS technique 'AI Agent Tool Invocation' (AML.T0053, 15 documented real cases) is the single most under-reported attack vector relative to its real-world prevalence, and the 70 stories on multi-agent risks in our feed are mostly about agent failures, not about attackers deliberately invoking agent tools to cause harm — meaning governance teams are preparing for accidents when they should be preparing for targeted abuse.

What would prove it wrong: If no such attack is disclosed by 2027-03, or if all disclosed agent-tool attacks are correctly attributed to external attackers in first reporting, then the framing gap is not as severe as claimed.

Read the post that made this call →
Open said it would be true by 2028-08

By 2028-08, when the EU AI Act Annex I embedded high-risk obligations apply, at least two major industrial AI deployments will face sudden regulatory halts due to unmeasured robustness failures that never appeared in prior incident feeds.

The argument it came from: The absence of reporting on environmental harm and system robustness is a structural failure of incident databases, not an indicator of zero risk.

What would prove it wrong: If all regulatory halts or enforcement actions under EU AI Act Annex I prior to August 2028 stem exclusively from risks already tracked in current top-five incident categories.

Read the post that made this call →
Open said it would be true by 2027-03

By 2027-03, public governance incident reports will continue to decline even as total internal AI incidents (as measured by corporate disclosures in securities filings) rise.

The argument it came from: The steep decline in governance-related AI incident reporting is not a sign of improvement but a leading indicator that organizations are shifting incident disclosure from public channels to private legal and compliance workflows ahead of enforceable deadlines.

What would prove it wrong: If public governance incident reports increase while compliance reports also rise, the shift hypothesis is wrong.

Read the post that made this call →
Open said it would be true by 2028-03

We will see an increase in incidents related to AI system failures or unsafe behavior that could have been prevented by adequate safety and robustness measures, despite the increased focus on privacy.

The argument it came from: The regulatory focus on privacy (88 stories in 45 days) is creating a 'privacy paradox' where organizations prioritize privacy controls over more fundamental safety and robustness measures, which are underreported (0 stories).

What would prove it wrong: If the number of reported incidents related to safety and robustness remains negligible or decreases, the prediction is wrong.

Read the post that made this call →
Open said it would be true by 2028-12

By December 2028, enterprises targeted by automated AI attacks will experience model evasion (AML.T0015) as their primary vector despite zero coverage in major trade outlets.

The argument it came from: The heavy concentration of media coverage in top-tier outlets on privacy and multi-agent risks creates a false sense of security that blinds leadership to unmentioned attack vectors like prompt crafting and model evasion.

What would prove it wrong: If enterprise AI security incidents continue to be dominated exclusively by privacy leaks and agent misbehavior rather than model evasion through 2028.

Read the post that made this call →
Open said it would be true by 2028-12

By 2028-06, we will see the first major financial fraud incident caused by compromised AI agent tool invocation that goes undetected for at least 30 days due to monitoring gaps.

The argument it came from: The near-total absence of reported incidents related to AI agent tool invocation (AML.T0053) indicates organizations are failing to monitor autonomous AI systems' external interactions, creating an undetected risk surface.

What would prove it wrong: If no such incident is reported in financial news outlets (e.g., WSJ, Bloomberg) or regulatory enforcement announcements by December 2028.

Read the post that made this call →
Open said it would be true by 2027-07

By July 2027, at least 40% of major enterprise AI security breaches will stem from robustness failures categorized as non-existent in current governance frameworks.

The argument it came from: Security intelligence briefings show 25 active ai-security articles while governance feeds report zero robustness stories, proving that technical exploitability is completely decoupled from enterprise risk registers.

What would prove it wrong: If enterprise AI incident disclosures attribute less than 10% of breaches to robustness or capability failures in published post-mortems through July 2027.

Read the post that made this call →
Open said it would be true by 2027-09

By 2027-09, a leaked internal document from a Fortune 500 company will show overreliance incidents were tracked internally but excluded from public incident reports, and this will trigger a class-action lawsuit.

The argument it came from: The 0 stories on 'Overreliance and unsafe use' (MIT AI Risk Repository 5.1) in 180 days is not a data gap but evidence that governance teams are actively suppressing this risk class because acknowledging it would invalidate their own AI literacy training programs, which are the primary compliance output they sell internally.

What would prove it wrong: If no such leak or lawsuit occurs by that date, and no regulator questions the absence of overreliance reporting, the prediction fails.

Read the post that made this call →
Open said it would be true by 2027-03

By early 2027, at least two major AI incidents will be revealed to have been reported as 'compliance issues' in 2026, with the underlying vulnerability still unpatched.

The argument it came from: The collapse in reported AI incident volume (from 450 to 344 stories) is not a decline in risk but a migration of incidents into the 'compliance' category, indicating that organizations are re-labeling failures to fit regulatory checklists rather than fixing underlying system flaws.

What would prove it wrong: If a retrospective audit of 2026 incident reports shows no systematic re-labeling from governance to compliance categories.

Read the post that made this call →
Open said it would be true by 2028-06

By 2028-06, at least 60% of AI incidents reported in major outlets will be classified as 'critical' severity, up from 25% in the last 45 days of 2026, while privacy incidents will remain the most frequently reported category.

The argument it came from: The recent surge in reported privacy incidents is masking a simultaneous increase in critical severity incidents across all categories, indicating a disproportionate focus on privacy that could leave organizations unprepared for more damaging events.

What would prove it wrong: If the proportion of critical incidents reported in major outlets remains below 40% of all reported incidents by 2028-06, or if privacy incidents drop to less than 40% of all reported incidents by then.

Read the post that made this call →
Open said it would be true by 2028-12

By December 2028, at least two major public safety or infrastructure outages will be directly attributed to unmitigated model robustness failures that were absent from the deploying company's risk register.

The argument it came from: Because trade press coverage ignores zero-story risk domains like robustness and overreliance, corporate risk registers will entirely omit these vectors until a catastrophic failure triggers retroactive litigation.

What would prove it wrong: Public post-mortem reports for major AI system failures in critical infrastructure showing that robustness and overreliance were explicitly rated and tracked prior to the incident.

Read the post that made this call →
Open said it would be true by 2028-02

By February 2028, at least two of the Big Four accounting firms will formally combine their AI governance and cybersecurity risk practices into a single unified service line.

The argument it came from: Security intelligence briefings warning that 'Hackers want your AI brains more than your bitcoin' will force the merger of traditional CISO vulnerability management with model risk governance by early 2028.

What would prove it wrong: Public corporate restructurings showing continuous separation of AI governance and cybersecurity advisory practices among all Big Four firms through February 2028.

Read the post that made this call →
Open said it would be true by 2028-03

By 2028-03, at least one major autonomous AI system will cause a significant operational failure due to overreliance, resulting in a $100M+ loss

The argument it came from: The complete absence of reported incidents on AI overreliance creates a blind spot that will lead to catastrophic operational failures in autonomous systems.

What would prove it wrong: If no major autonomous AI system failure exceeding $100M occurs due to overreliance by this date

Read the post that made this call →
Open said it would be true by 2028-01

Publicly disclosed internal AI algorithm failures by Fortune 500 companies operating in California will decline by at least 25% in the twelve months following the January 2027 CPPA effective date.

The argument it came from: The impending enforcement of the California CPPA ADMT rules in early 2027 will cause enterprises to aggressively suppress internal AI incident logging to avoid creating discoverable compliance evidence.

What would prove it wrong: If public disclosures of internal AI algorithm failures by Fortune 500 companies increase or remain flat through January 2028.

Read the post that made this call →

What the agents changed with no human

When the official text behind a regulation moves, an agent reads it and drafts the new status. It only publishes if three independent readings agree on the confidence, on the status category, and on whether the change is material at all. Any dissent and nothing is published. Every change it did publish is below, with the before and after.

No agent has published a change on its own yet. This stays empty until three independent readings agree on one — which is the bar, and most candidates never clear it.

What we threw away

The promoter has judged 1379 candidates from the live news feed and kept 0. It last judged something 3h ago, with 2 waiting.

100% of everything it looked at was rejected. That is the intended outcome — most AI news is commentary, not an incident — and the reasons are below.

  • 1221 not actually an incident
  • 158 too thin to write up

Links we publish that are broken

Every regulation page links its official text. A watcher re-fetches all of them on a six-hour cycle, so a dead one is our broken promise and gets listed here rather than quietly sitting on the page. A further 4 hosts refuse automated checks entirely; those are verified by hand and are not counted as broken.

dead link last checked 9h ago

Illinois HB 3773 — AI in Employment (IHRA amendment)

https://www.ilga.gov/legislation/PublicActs/View/103-0804

unable to verify the first certificate

The page that publishes this link →
dead link last checked 15h ago

Biometric Information Privacy Act (BIPA)

https://www.ilga.gov/Legislation/ILCS/Articles?ActID=3004&ChapterID=57

unable to verify the first certificate

The page that publishes this link →

How to check any of this

  • Every regulation page carries the date its own official source was last read, and warns you when it has moved since we wrote the summary — start at the regulation reference.
  • Every curated incident says what it was, what failed, and the control that would have prevented it — the incident radar.
  • Every prediction above links the post that made it. The posts are dated and were not edited afterwards.
  • Every dated obligation we publish has its own page and a calendar feed you can subscribe to and check against the official source yourself.
  • Methodology & Scoring gives the actual formula behind every score and classification on this site, so you can recompute one rather than trust it.