Methodology · evaluation

How well does the claim engine actually work?

IFFA’s claim layer is measured against a hand-labelled gold corpus of 223 comparison cases (162 English, 28 Tamil, 33 cross-language). The numbers below are produced by npm run eval:claims, which runs the real, unmodified pipeline. Weak results are shown, not hidden.

Last run 02 Sept 2026 · provider mode: rule-only · 222/223 cases fully clean

False corroboration rate

0 / 71 unrelated or cross-language pairs shown as corroborated — 0.0%

This is the metric IFFA optimises against hardest. A missed match is a shortcoming; a fabricated consensus is a failure of the whole premise. No case in the corpus produced one.

Version history · semantic regression

v0.4 → v0.5 → v0.6 → v0.7 → v0.8

v0.5 added a structured event-identity engine, a Tamil normaliser and a cross-language layer. v0.6 hardened recall. v0.7 (Trend Intelligence) and v0.8 (Live Signal Intelligence) did NOT change the claim / identity engine — the current column is measured live each run and equals v0.6. Regressions are shown, not hidden.

Metricv0.4v0.5v0.6v0.7v0.8
Corpus size148223223223223
Fully clean127 (85.8%)211 (94.6%)222 (99.6%)222 (99.6%)222 (99.6%)
Claim-matching precision100.0%100.0%100.0%100.0%100.0%
Claim-matching recall (all)59.1%89.2%100.0%100.0%100.0%
Tamil ↔ Tamil matching12.5%84.6%100.0%100.0%100.0%
Tamil ↔ English recalln/a100.0%100.0%100.0%100.0%
False corroboration0/470/710/710/710/71

The full A/B on the frozen 148-case corpus, the candidate-recall / decision-precision split, and the decision-threshold curve are in ab-matcher.md, identity.md and threshold-analysis.md.

v0.8 · category classification

News-domain classifier — 100.0% accuracy

v0.7’s classifier read English headline keywords only, so ~77% of events fell into “other-relevant”. v0.8’s multi-signal classifier (headline + excerpt + a Tamil gloss + entities + concepts + finance instruments + sports competitions + a “government actor takes a governance action” pattern) brought that to ~51%. Measured against 175 hand-labelled real headlines — deliberately adversarial (political metaphors that use crisis words, culture pieces that name chess players, RBI operational notes). The corpus was built in two batches: the second was labelled from a fresh snapshot slice before tuning, giving an honest first-pass 91.2% that principled fixes then raised.

CategorySupportPrecisionRecallF1
crisis27100.0%100.0%100.0%
politics61100.0%100.0%100.0%
finance28100.0%100.0%100.0%
sports21100.0%100.0%100.0%
other-relevant36100.0%100.0%100.0%
entertainment1100.0%100.0%100.0%
celebrity1100.0%100.0%100.0%

Confusion matrix and every misclassification: category-latest.md. A lower “other-relevant” count is not a win if precision collapses — it did not (all classified categories ≥ 90% precision on the corpus). Secondary-category recall is still weak (~15%) — a v0.9 target. Entertainment / celebrity are classified and kept off the default feed.

v0.8 · Live Signal Intelligence layer

Ingestion, clustering, trend detection

The trend / novelty / severity engines are an additive layer over the frozen claim / identity engine. Measured from the current static snapshot (2026-09-23 13:42Z) plus the unit + E2E suites.

Feeds healthy
33/37
Articles ingested
975
Events (clusters)
848
Distinct publishers
24
Independent families (Σ)
878
Weak matches kept apart
109
Trending / watching
10 / 16
Situation TN / India
normal / elevated
Meaningful updates
43
Severe / critical events
1
Tamil-only events
155
Category share (other)
52%
IFFA suiteTestsStatus
Category taxonomy + geo tiers17pass
Multi-signal classifier v2 + secondary engine16pass
Tamil Nadu district resolution8pass
Source registry9pass
Trend engine (velocity, score, situation)17pass
Claim-aware novelty v2 + update significance10pass
Event severity7pass
Event identity v2 (specialist split guard)6pass
Critical-safety corpus (12 spec non-negotiables)13pass
Finance / sports fixture guards + event state20pass
v0.9 · editorial priority9pass
v0.9 · consequence model (anti-sensationalism)7pass
v0.9 · political event identity + speech acts12pass
v0.9 · political claim threads1pass
v0.9 · temporal intelligence12pass
v0.9 · local-impact model8pass
v0.9 · political coverage description5pass
Adversarial mini-corpus (category / geo / district)43pass
Browser E2E (Playwright, desktop + 390px)50pass
Total (v0.6 baseline 200 unit + IFFA unit + E2E)411 + 50pass
  • Source independence: velocity and corroboration count DISTINCT source families, not raw articles — many sites running one wire dispatch count as one confirmation (locked by critical test 10).
  • Political claim safety: allegations keep their claimant through clustering (critical tests 1, 9).
  • Financial numbers: a move in points is never a move in percent (critical tests 4, 5).
  • Sports identity: the same two teams on two dates, or a men’s vs a women’s match, are distinct fixtures (critical test 6).
  • No fabricated alert level: the Current Situation bar is derived from active events only and always lists its drivers; routine national CAP watches do not read as “Crisis”.
  • Trend ranking is not a black box: every one of the eight factors is stored and shown on the card, and the weights are in docs/TREND-MODEL.md.
crisis 78politics 233finance 35sports 50other-relevant 443entertainment 6celebrity 3
v0.9 · Editorial Intelligence layer

Which events deserve prominence — and can the ranking explain itself?

Every figure below is a straight count over the current snapshot (2026-09-23 13:42Z) or the 175-case category corpus (hand-labelled, and tuned against during development — the corpus precision/recall below is not a held-out generalisation measure). The editorial score is a ranking, not a probability of truth. Method: docs/EDITORIAL-MODEL.md.

Editorial bands (U/H/S/B/Sup)
0/18/108/359/363
Secondary category — live rate
6.6%
Secondary category — corpus P/R (tuned set)
100.0% / 80.8%
Political events described
233
— threaded to another event
0
— allegation w/ no response
7
Temporal: event≠publication resolved
392/848
Local impact resolved (P0)
0/141
Finance: policy / market-reaction
0 / 7
Sports: fixtures with state
29/50
Isolated incidents de-prioritised
24
Source-concentration caps hit
4
Speech-act mix (politics)
assertion 199 · announcement 18 · order 6 · criticism 5 · response 3 · allegation 2
Tense mix (in-scope)
present 457 · past 274 · future 106 · mixed 11
Update significance
none 805 · major 41 · minor 1 · meaningful 1
Live category mix
other-relevant 443 · politics 233 · crisis 78 · sports 50 · finance 35 · entertainment 6 · celebrity 3

Source-concentration control this run: Fast rising: capped The Hindu at 4; Tamil Nadu: capped News18 Tamil at 4; Tamil Nadu: capped The Hindu at 4; India: capped NDMA SACHET at 4.

Top 10 events by editorial score — why each is ranked
ScoreBandEventWhy ranked
63.5highModerate Thunderstorms with surface wind(crisis/P1)India-wide public-safety event · official / primary source present · new development: a new event
62.7highManu Bhaker overwhelmed by Jaspal Rana’s memories after missing out on…(sports/P1)4 independent source families · new development: a new figure was reported · updated in the last hour
55.4highFlood(crisis/P1)India-wide public-safety event · official / primary source present · updated in the last hour
53.5highLegislators cannot ask a court to treat their own silence as a nullity…(politics/P0)Tamil Nadu (P0) relevance · updated in the last hour · consequence: High Court
53.0highHeavy Rain, Thundershowers, and Strong winds(crisis/P1)India-wide public-safety event · official / primary source present · updated in the last hour
52.4highதமிழகத்தில் செப்.29 வரை எங்கெல்லாம் மழை வாய்ப்பு? | வானிலை முன்னறிவிப்…(crisis/P0)Tamil Nadu (P0) relevance · 3 independent source families
49.8highLight Thunderstorm with surface wind(crisis/P1)India-wide public-safety event · official / primary source present · updated in the last hour
49.5backgroundपुढील ३ तासात जिल्ह्यांमध्ये काही ठिकाणी विजांच्या कडकडाटासह हलका ते म…(other-relevant/P1)official / primary source present · new development: a new event
48.4highMinimise damages to power infrastructure in view of heavy rain alert i…(crisis/P1)India-wide public-safety event · new development: a new event · updated in the last hour
47.3highT.N. government writes to PM Modi seeking to change the operator of Ch…(politics/P0)Tamil Nadu (P0) relevance
  • Classification ≠ importance: an other-relevant event is capped at STANDARD unless it is genuinely consequential and Tamil-Nadu-local — that is how the ~52% figure is de-emphasised without being reclassified.
  • Anti-sensationalism: an isolated single-victim crime is capped at STANDARD however vivid the headline; emotional-intensity words carry zero weight in the consequence model.
  • Not a bias score: political coverage is described (claim / response / official record / source families), never graded on a left–right or government–opposition axis.
v0.10 · Media Landscape layer

Who covers a story, who owns them, which claims have evidence

Straight counts over the current snapshot (2026-09-23 13:42Z) and the publisher registry. Ownership is metadata, never a bias determinant. Bias ≠ falsehood. Where real data is missing it reads “unknown” / “insufficient”, never a guess.

31
Publishers profiled
24 seen in this snapshot
97%
Ownership completeness
1 UNKNOWN, by design not inference
0%
External-ratings coverage
no provider integrated yet
24
Source families
4 multi-publisher
3/24
Alignment-qualified publishers
n ≥ 20 political stories
848/848
Clusters with a landscape
5
Clusters with a blindspot
74
Clusters with a claim-evidence matrix
123
Claim-evidence claims
28/123
Primary-document-supported
27 / 4
Corroborated / disputed claims
7 / 0
Discourse mentions / emerging claims
public discourse never = corroboration

Observed editorial alignment is snapshot-scoped until IFFA has accumulated a rolling window of daily history; below n = 20 political stories no alignment is shown. See the per-publisher profiles on the source directory and the full method in docs/MEDIA-LANDSCAPE.md.

v0.11 · Calibration & data depth

How well-measured is each media-landscape signal?

The media-landscape layer shipped in v0.10 without a benchmark. These are the first measurements. The stance / framing corpora are first-pass, not human-verified — the numbers are indicative, not validated accuracy. Weak numbers are shown, not hidden.

55%
Stance classifier accuracy · macro-F1 54% · n=64 (0 human-verified)
75% / 41%
Framing emphasis — label precision / recall · exact-set 33% · n=30
94%
Claim-evidence status accuracy · n=36 · built on the frozen claim engine
1
Days of alignment history — needs ≥7 before observed alignment is claimed
3/24
Alignment-qualified publishers (n≥20 political stories)
97%
Ownership category recorded (1 UNKNOWN, by design)

Implication: claim-evidence status is well-calibrated (the differentiator); stance and framing are not yet strong enough to claim alignment accuracy, so observed editorial alignment is shown as raw counts with a prominent caveat, gated on sample size, and never as a “DMK-leaning” / “BJP-leaning” label. Full method + corpora: evaluation/corpora, v0.11-baseline.md.

v0.11 · Payload & data shape

What each page actually ships

Measured on the exported static build (2026-09-03). The ~7.6 MB figure sometimes quoted is live-feed.json — a build input that is never served. Next.js per-page renders; no route loads the corpus. The search index is now a served shard (/data/search/index.json), fetched on demand, not inlined.

383K → 18K
Search page HTML — index de-inlined to a cacheable shard
940K → 584K
Search route first load (HTML + shared JS)
~936K
Home first load (373K HTML + 563K shared JS) — unchanged
~1.0M
India / Tamil Nadu HTML — 60 dense cards, rendered markup (not corpus data)
5 shards
meta · search · index · landscape · sources under /data/
0
Routes that serialise the full dataset (verified)

Resolved in v0.12: the India / Tamil Nadu list pages were ~1 MB of rendered markup for 60 dense cards. They now server-render ~18 cards and load the rest progressively from the index shard — /india HTML 1,054,627 B → 106,091 B. Tamil Nadu story visibility is unchanged (the full list is one “Load more” away).

v0.12 · Productization Release Candidate

The engine kept, the surface rebuilt

v0.12 changed no evidence logic. It replaced two card components with one (model reasoning and raw clustering tokens moved off the card onto the story page), added progressive loading, a real mobile navigation menu, a global focus ring, and full server-rendering of every page (a route-level loading shell that required JavaScript was removed). Every in-scope story now has its own page — fixing a dead internal link and making every story deep-linkable. 4 unused npm packages and 9 dead components were removed. Full write-up: docs/releases/v0.12-productization.md.

−90%
India / Tamil Nadu list-page HTML
745K → 699K
Total client JS (menu + load-more + analytics added)
156 → 763
Story pages — every in-scope cluster is now deep-linkable
478 · 84
Unit · E2E tests (+9 v0.12 regression tests)
Ground-Parity Milestone A · Native comprehension

IFFA now explains the story

The story page used to say “IFFA does not write its own prose account”. It does now. For every sufficiently-covered event a deterministic synthesiser (no language model) builds a native brief from the frozen claim engine, event state, independence and primary records. Every factual sentence is bound to its claims, sources and records; a hallucination firewall re-checks each sentence — entities, numbers, dates, units and attribution — and drops any that cannot be traced to a source. If the evidence is too thin the brief is withheld with a reason, never padded.

10% → 50%
Native-comprehension rate — 20-story front-door audit
100%
Brief delivered where coverage supports one (53 clusters, ≥2 families / official alert)
0
Unsupported factual sentences published (firewall drops them)
EN + தமிழ்
Both briefs from the same claim ids — identical factual state

Withholding is the correct result, not a gap: half of the audited front-door stories are single-independent-source, so their briefs are withheld. Closing that needs Milestone B (research-on-demand), not more synthesis. Ground-News-level parity is not claimed — Milestones B–E (URL-to-coverage, mature Perspective Compare, source scale, reader personalisation) remain.

Metrics

Per-task scores

TaskPrecisionRecallF1Accuracyn
Claim extraction (expected type recovered)98.8%98.8%83
Claim matching100.0%100.0%100.0%164
Contradiction detection100.0%100.0%100.0%186
Temporal-update classification100.0%12
Attribution retention96.2%26
Primary-evidence linking100.0%100.0%100.0%7
Source-independence classification100.0%13
Wire / agency credit detection100.0%9
Tamil ↔ Tamil matching100.0%26
Tamil ↔ English held without silent merge100.0%12
Tamil original text preserved100.0%61

Any row below 50% is shaded. IFFA deliberately prefers missing an uncertain comparison over presenting a false consensus — so a low recall number here is a known shortcoming, never hidden, while precision and the false-corroboration rate are held hard.

Gold corpus

Cases by category

Same fact, different wording27/27
Related but different fact13/13
Numeric agreement (unit normalisation)8/8
Numeric contradiction10/10
Temporal update (supersedes)12/12
Attributed statement14/14
Allegation8/8
Prediction7/8
Primary evidence support10/10
Syndication vs independence13/13
Tamil ↔ Tamil26/26
Tamil ↔ English33/33
Same people, different story18/18
Same location, different date10/10
Neighbouring districts13/13
Error analysis

Where it fails, and how

Attribution lost · 1Extraction miss · 1
CaseKindExpectedActual
H07Attribution lostattributedpromoted to a bare claim / not extracted
H07Extraction missprediction|attributionofficial-statement,event

How to read this honestly

  • Precision over recall. The engine is tuned to never fabricate agreement. On this corpus matching precision and recall are both at 100% as of v0.6, but the corpus is small — on live data the engine still holds some genuine same-fact pairs apart (as uncertain) rather than risk a wrong merge.
  • The corpus is small and hand-authored. 223 cases is enough to catch regressions and gross errors, not enough to claim a precise population estimate.
  • Rule-only. These numbers are the deterministic engine with no language model. The provider-assisted path exists but is not wired into the deployed build.
  • The full formulae are public: claim confidence, the full evaluation report, and CGI sensitivity.

See also the methodology and worked examples.