Methodology · evaluation
How well does the claim engine actually work?
IFFA’s claim layer is measured against a hand-labelled gold corpus of 223 comparison cases (162 English, 28 Tamil, 33 cross-language). The numbers below are produced by npm run eval:claims, which runs the real, unmodified pipeline. Weak results are shown, not hidden.
Last run 02 Sept 2026 · provider mode: rule-only · 222/223 cases fully clean
0 / 71 unrelated or cross-language pairs shown as corroborated — 0.0%
This is the metric IFFA optimises against hardest. A missed match is a shortcoming; a fabricated consensus is a failure of the whole premise. No case in the corpus produced one.
v0.4 → v0.5 → v0.6 → v0.7 → v0.8
v0.5 added a structured event-identity engine, a Tamil normaliser and a cross-language layer. v0.6 hardened recall. v0.7 (Trend Intelligence) and v0.8 (Live Signal Intelligence) did NOT change the claim / identity engine — the current column is measured live each run and equals v0.6. Regressions are shown, not hidden.
| Metric | v0.4 | v0.5 | v0.6 | v0.7 | v0.8 |
|---|---|---|---|---|---|
| Corpus size | 148 | 223 | 223 | 223 | 223 |
| Fully clean | 127 (85.8%) | 211 (94.6%) | 222 (99.6%) | 222 (99.6%) | 222 (99.6%) |
| Claim-matching precision | 100.0% | 100.0% | 100.0% | 100.0% | 100.0% |
| Claim-matching recall (all) | 59.1% | 89.2% | 100.0% | 100.0% | 100.0% |
| Tamil ↔ Tamil matching | 12.5% | 84.6% | 100.0% | 100.0% | 100.0% |
| Tamil ↔ English recall | n/a | 100.0% | 100.0% | 100.0% | 100.0% |
| False corroboration | 0/47 | 0/71 | 0/71 | 0/71 | 0/71 |
The full A/B on the frozen 148-case corpus, the candidate-recall / decision-precision split, and the decision-threshold curve are in ab-matcher.md, identity.md and threshold-analysis.md.
News-domain classifier — 100.0% accuracy
v0.7’s classifier read English headline keywords only, so ~77% of events fell into “other-relevant”. v0.8’s multi-signal classifier (headline + excerpt + a Tamil gloss + entities + concepts + finance instruments + sports competitions + a “government actor takes a governance action” pattern) brought that to ~51%. Measured against 175 hand-labelled real headlines — deliberately adversarial (political metaphors that use crisis words, culture pieces that name chess players, RBI operational notes). The corpus was built in two batches: the second was labelled from a fresh snapshot slice before tuning, giving an honest first-pass 91.2% that principled fixes then raised.
| Category | Support | Precision | Recall | F1 |
|---|---|---|---|---|
| crisis | 27 | 100.0% | 100.0% | 100.0% |
| politics | 61 | 100.0% | 100.0% | 100.0% |
| finance | 28 | 100.0% | 100.0% | 100.0% |
| sports | 21 | 100.0% | 100.0% | 100.0% |
| other-relevant | 36 | 100.0% | 100.0% | 100.0% |
| entertainment | 1 | 100.0% | 100.0% | 100.0% |
| celebrity | 1 | 100.0% | 100.0% | 100.0% |
Confusion matrix and every misclassification: category-latest.md. A lower “other-relevant” count is not a win if precision collapses — it did not (all classified categories ≥ 90% precision on the corpus). Secondary-category recall is still weak (~15%) — a v0.9 target. Entertainment / celebrity are classified and kept off the default feed.
Ingestion, clustering, trend detection
The trend / novelty / severity engines are an additive layer over the frozen claim / identity engine. Measured from the current static snapshot (2026-09-23 13:42Z) plus the unit + E2E suites.
| IFFA suite | Tests | Status |
|---|---|---|
| Category taxonomy + geo tiers | 17 | pass |
| Multi-signal classifier v2 + secondary engine | 16 | pass |
| Tamil Nadu district resolution | 8 | pass |
| Source registry | 9 | pass |
| Trend engine (velocity, score, situation) | 17 | pass |
| Claim-aware novelty v2 + update significance | 10 | pass |
| Event severity | 7 | pass |
| Event identity v2 (specialist split guard) | 6 | pass |
| Critical-safety corpus (12 spec non-negotiables) | 13 | pass |
| Finance / sports fixture guards + event state | 20 | pass |
| v0.9 · editorial priority | 9 | pass |
| v0.9 · consequence model (anti-sensationalism) | 7 | pass |
| v0.9 · political event identity + speech acts | 12 | pass |
| v0.9 · political claim threads | 1 | pass |
| v0.9 · temporal intelligence | 12 | pass |
| v0.9 · local-impact model | 8 | pass |
| v0.9 · political coverage description | 5 | pass |
| Adversarial mini-corpus (category / geo / district) | 43 | pass |
| Browser E2E (Playwright, desktop + 390px) | 50 | pass |
| Total (v0.6 baseline 200 unit + IFFA unit + E2E) | 411 + 50 | pass |
- Source independence: velocity and corroboration count DISTINCT source families, not raw articles — many sites running one wire dispatch count as one confirmation (locked by critical test 10).
- Political claim safety: allegations keep their claimant through clustering (critical tests 1, 9).
- Financial numbers: a move in points is never a move in percent (critical tests 4, 5).
- Sports identity: the same two teams on two dates, or a men’s vs a women’s match, are distinct fixtures (critical test 6).
- No fabricated alert level: the Current Situation bar is derived from active events only and always lists its drivers; routine national CAP watches do not read as “Crisis”.
- Trend ranking is not a black box: every one of the eight factors is stored and shown on the card, and the weights are in docs/TREND-MODEL.md.
Which events deserve prominence — and can the ranking explain itself?
Every figure below is a straight count over the current snapshot (2026-09-23 13:42Z) or the 175-case category corpus (hand-labelled, and tuned against during development — the corpus precision/recall below is not a held-out generalisation measure). The editorial score is a ranking, not a probability of truth. Method: docs/EDITORIAL-MODEL.md.
Source-concentration control this run: Fast rising: capped The Hindu at 4; Tamil Nadu: capped News18 Tamil at 4; Tamil Nadu: capped The Hindu at 4; India: capped NDMA SACHET at 4.
| Score | Band | Event | Why ranked |
|---|---|---|---|
| 63.5 | high | Moderate Thunderstorms with surface wind(crisis/P1) | India-wide public-safety event · official / primary source present · new development: a new event |
| 62.7 | high | Manu Bhaker overwhelmed by Jaspal Rana’s memories after missing out on…(sports/P1) | 4 independent source families · new development: a new figure was reported · updated in the last hour |
| 55.4 | high | Flood(crisis/P1) | India-wide public-safety event · official / primary source present · updated in the last hour |
| 53.5 | high | Legislators cannot ask a court to treat their own silence as a nullity…(politics/P0) | Tamil Nadu (P0) relevance · updated in the last hour · consequence: High Court |
| 53.0 | high | Heavy Rain, Thundershowers, and Strong winds(crisis/P1) | India-wide public-safety event · official / primary source present · updated in the last hour |
| 52.4 | high | தமிழகத்தில் செப்.29 வரை எங்கெல்லாம் மழை வாய்ப்பு? | வானிலை முன்னறிவிப்…(crisis/P0) | Tamil Nadu (P0) relevance · 3 independent source families |
| 49.8 | high | Light Thunderstorm with surface wind(crisis/P1) | India-wide public-safety event · official / primary source present · updated in the last hour |
| 49.5 | background | पुढील ३ तासात जिल्ह्यांमध्ये काही ठिकाणी विजांच्या कडकडाटासह हलका ते म…(other-relevant/P1) | official / primary source present · new development: a new event |
| 48.4 | high | Minimise damages to power infrastructure in view of heavy rain alert i…(crisis/P1) | India-wide public-safety event · new development: a new event · updated in the last hour |
| 47.3 | high | T.N. government writes to PM Modi seeking to change the operator of Ch…(politics/P0) | Tamil Nadu (P0) relevance |
- Classification ≠ importance: an
other-relevantevent is capped at STANDARD unless it is genuinely consequential and Tamil-Nadu-local — that is how the ~52% figure is de-emphasised without being reclassified. - Anti-sensationalism: an isolated single-victim crime is capped at STANDARD however vivid the headline; emotional-intensity words carry zero weight in the consequence model.
- Not a bias score: political coverage is described (claim / response / official record / source families), never graded on a left–right or government–opposition axis.
Who covers a story, who owns them, which claims have evidence
Straight counts over the current snapshot (2026-09-23 13:42Z) and the publisher registry. Ownership is metadata, never a bias determinant. Bias ≠ falsehood. Where real data is missing it reads “unknown” / “insufficient”, never a guess.
Observed editorial alignment is snapshot-scoped until IFFA has accumulated a rolling window of daily history; below n = 20 political stories no alignment is shown. See the per-publisher profiles on the source directory and the full method in docs/MEDIA-LANDSCAPE.md.
How well-measured is each media-landscape signal?
The media-landscape layer shipped in v0.10 without a benchmark. These are the first measurements. The stance / framing corpora are first-pass, not human-verified — the numbers are indicative, not validated accuracy. Weak numbers are shown, not hidden.
Implication: claim-evidence status is well-calibrated (the differentiator); stance and framing are not yet strong enough to claim alignment accuracy, so observed editorial alignment is shown as raw counts with a prominent caveat, gated on sample size, and never as a “DMK-leaning” / “BJP-leaning” label. Full method + corpora: evaluation/corpora, v0.11-baseline.md.
What each page actually ships
Measured on the exported static build (2026-09-03). The ~7.6 MB figure sometimes quoted is live-feed.json — a build input that is never served. Next.js per-page renders; no route loads the corpus. The search index is now a served shard (/data/search/index.json), fetched on demand, not inlined.
meta · search · index · landscape · sources under /data/Resolved in v0.12: the India / Tamil Nadu list pages were ~1 MB of rendered markup for 60 dense cards. They now server-render ~18 cards and load the rest progressively from the index shard — /india HTML 1,054,627 B → 106,091 B. Tamil Nadu story visibility is unchanged (the full list is one “Load more” away).
The engine kept, the surface rebuilt
v0.12 changed no evidence logic. It replaced two card components with one (model reasoning and raw clustering tokens moved off the card onto the story page), added progressive loading, a real mobile navigation menu, a global focus ring, and full server-rendering of every page (a route-level loading shell that required JavaScript was removed). Every in-scope story now has its own page — fixing a dead internal link and making every story deep-linkable. 4 unused npm packages and 9 dead components were removed. Full write-up: docs/releases/v0.12-productization.md.
IFFA now explains the story
The story page used to say “IFFA does not write its own prose account”. It does now. For every sufficiently-covered event a deterministic synthesiser (no language model) builds a native brief from the frozen claim engine, event state, independence and primary records. Every factual sentence is bound to its claims, sources and records; a hallucination firewall re-checks each sentence — entities, numbers, dates, units and attribution — and drops any that cannot be traced to a source. If the evidence is too thin the brief is withheld with a reason, never padded.
Withholding is the correct result, not a gap: half of the audited front-door stories are single-independent-source, so their briefs are withheld. Closing that needs Milestone B (research-on-demand), not more synthesis. Ground-News-level parity is not claimed — Milestones B–E (URL-to-coverage, mature Perspective Compare, source scale, reader personalisation) remain.
Per-task scores
| Task | Precision | Recall | F1 | Accuracy | n |
|---|---|---|---|---|---|
| Claim extraction (expected type recovered) | — | 98.8% | — | 98.8% | 83 |
| Claim matching | 100.0% | 100.0% | 100.0% | — | 164 |
| Contradiction detection | 100.0% | 100.0% | 100.0% | — | 186 |
| Temporal-update classification | — | — | — | 100.0% | 12 |
| Attribution retention | — | — | — | 96.2% | 26 |
| Primary-evidence linking | 100.0% | 100.0% | 100.0% | — | 7 |
| Source-independence classification | — | — | — | 100.0% | 13 |
| Wire / agency credit detection | — | — | — | 100.0% | 9 |
| Tamil ↔ Tamil matching | — | — | — | 100.0% | 26 |
| Tamil ↔ English held without silent merge | — | — | — | 100.0% | 12 |
| Tamil original text preserved | — | — | — | 100.0% | 61 |
Any row below 50% is shaded. IFFA deliberately prefers missing an uncertain comparison over presenting a false consensus — so a low recall number here is a known shortcoming, never hidden, while precision and the false-corroboration rate are held hard.
Cases by category
Where it fails, and how
| Case | Kind | Expected | Actual |
|---|---|---|---|
| H07 | Attribution lost | attributed | promoted to a bare claim / not extracted |
| H07 | Extraction miss | prediction|attribution | official-statement,event |
How to read this honestly
- Precision over recall. The engine is tuned to never fabricate agreement. On this corpus matching precision and recall are both at 100% as of v0.6, but the corpus is small — on live data the engine still holds some genuine same-fact pairs apart (as uncertain) rather than risk a wrong merge.
- The corpus is small and hand-authored. 223 cases is enough to catch regressions and gross errors, not enough to claim a precise population estimate.
- Rule-only. These numbers are the deterministic engine with no language model. The provider-assisted path exists but is not wired into the deployed build.
- The full formulae are public: claim confidence, the full evaluation report, and CGI sensitivity.
See also the methodology and worked examples.