About & Methodology

A crisis-first, evidence-oriented news comparison platform for Tamil Nadu and India.

Information abundance does not automatically create understanding. IFFA groups public alerts and reporting around the same event so a reader can see what the official alert says, which independent sources confirm or contextualise it, what remains uncertain, and when the information was last refreshed.

Media landscape (v0.10)

For every story IFFA shows who is reporting it, who is not, who owns those sources, how their headlines and framing differ, which claims the reporting agrees on, which are disputed, and which have primary-document evidence. Ownership is provenance-backed metadata — it never determines a publisher’s alignment or reliability, and UNKNOWN is used wherever it is unverified. There is no single bias score; observed editorial alignment is corpus-derived, entity-specific (never a US left/right axis), and is withheld below a documented sample size. Bias is not falsehood, coverage asymmetry is not falsehood, and forum consensus is not evidence — see docs/MEDIA-LANDSCAPE.md and the methodology hub.

Scope

P0: Tamil Nadu (district-level). P1: India national. P2: events abroad only when they materially affect Tamil Nadu, India, Indian citizens, the economy, markets or foreign policy, or are a major global crisis. Category priority is Crisis → Politics → Finance → Sports; entertainment and celebrity stories are classified but kept out of the default feed (disabled, not deleted).

Category, novelty & severity (v0.8)

Every event is filed into a news domain — crisis, politics, finance, sports, or general — by a deterministic classifier that reads the headline, the excerpt, a Tamil-to-English gloss, the extracted entities, and signals like a financial instrument or a sports competition. It reports a confidence class(strong / moderate / weak / unknown), never a fake probability, and the quality dashboard shows its precision and recall against a hand-labelled corpus. A separate novelty check compares each new report against what the event already established — a headline rewrite is not a meaningful update; a corrected death toll or a first official confirmation is. Crisis events also carry an event severity (informational → watch → significant → severe → critical) derived from casualty counts and confirmed impact — this describes how bad the event is, not whether the reports are true; provenance is tracked separately.

Editorial priority (v0.9)

IFFA ingests broadly but displays selectively. On top of the trend score sits an editorial priority — a ranking score that decides how much prominence an event gets on the home page. It is a weighted mean of eight interpretable factors (geographic relevance, consequence, information gain, category, corroboration, meaningful recency, local impact, velocity), minus named penalties for churn, staleness, syndication and thin evidence. Every factor and penalty is shown on the card and on the quality dashboard. The score is a ranking, not a probability of truth. A gruesome single-victim crime is capped below the front strip however vivid its wording; general-interest news is de-emphasised editorially rather than reclassified; and political coverage is described (claim / response / official record / source families), never scored on a left–right axis. Full method: docs/EDITORIAL-MODEL.md.

Trend ranking (v0.7)

IFFA ranks events, not articles, and by what is changing, not publication count. The trend score is a weighted geometric mean of eight factors — recency, publication velocity across independent newsrooms, source diversity, geographic relevance, category, consequence, novelty, and corroboration. Every factor is shown on the card and the weights are public. Velocity counts independent source families, so many sites reprinting one wire dispatch count as a single confirmation. “Watching” holds stories that matter but lack the independent evidence to be called trending — a single local report of a bridge collapse never becomes a “confirmed crisis”.

The pipeline

  1. 01 FetchConfigured RSS / Atom / CAP feeds are fetched with a 15-second timeout and an identifying user agent. One feed failing never aborts the run.
  2. 02 Normalise & sanitiseEvery externally sourced string is stripped of markup, decoded, cleaned of control characters and length-clamped. Items without a valid source URL or a parseable date are rejected.
  3. 03 DeduplicateBy canonical URL and by normalised-headline tokens within a source.
  4. 04 Geo-classifyA rule-based Tamil Nadu dictionary (38 districts + state terms + Tamil-script tokens) assigns scope: tamil-nadu / india / india-relevant / excluded. Every classification carries the terms that matched and a reason.
  5. 05 Crisis-classifyDeterministic matchers detect priority incident types (cyclone, flood, dam warning, coastal warning, earthquake, landslide, heatwave, industrial accident, and more). CAP disaster-type from an official alert is trusted directly.
  6. 06 RankA reproducible 0–100 priority from: official-alert status, crisis-type weight, CAP severity/urgency/certainty (preserved verbatim), Tamil Nadu match, affected district count, recency, corroborating source count and primary documentation. Expired and all-clear alerts are scored down and kept out of the active banner.
  7. 07 ClusterTwo reports join a cluster only when time window, event type and geography all align and headline tokens overlap — or an official alert's key terms are contained in a report about the same districts. Unrelated events are not merged just because both mention rain.

Labels

This edition does not apply Left / Center / Right political-orientation ratings to Indian publications. Each report instead carries an evidence role:

  • Official alert
  • Primary document
  • Government statement
  • On-ground report
  • Independent report
  • Expert analysis
  • Developing / unverified

Reliability is expressed as evidence status — how well corroborated a claim is, not an ideological judgement:

  • Official primary source
  • Independently corroborated
  • Single-source report
  • Developing
  • Disputed
  • Unverified

Grounded claims

On a multi-source event, IFFA breaks the coverage into structured claims and classifies each one: corroborated (more than one independent source group), single source, attributed (something a named speaker said, alleged, expects or warned — kept as the speaker’s claim, never promoted to a bare fact unless separate evidence supports it), disputed, or outdated. Each claim carries a documented confidence score — the formula is public. Publication count is not corroboration count: a dedicated independence engine classifies every pair of reports as independent, syndicated, or unclear — several outlets running one PTI dispatch count as a single confirmation, and “unclear” never counts as independent. The extraction is deterministic and rule-based — no language model in the deployed build — and wording may not be exact, so the original source text is always linked. The Common Ground Index is experimental and describes the state of the reporting, not a verdict on the event.

The claim engine is measured against a hand-labelled gold corpus. See the claim-quality dashboard for extraction, matching, contradiction and attribution scores — including the ones that are still weak.

Copyright & provenance

  • IFFA stores only the headline, source name, canonical URL, publication timestamp, a feed-provided short excerpt, and structured alert metadata.
  • It never copies full article bodies, removes attribution, or circumvents a paywall.
  • Every item links to the original publisher for the full report.
  • IFFA summaries are never presented as the publisher’s own wording.

Limitations

  • IFFA does not claim algorithmic neutrality. Clustering, geo-classification and claim extraction are rule-based and can err.
  • Claims are extracted by deterministic rules from headlines and short excerpts. A structured event-identity engine now recovers every same-fact pair in the labelled corpus, but on live data the engine still holds some genuine matches apart — as uncertain rather than risk a wrong merge. It never merges a pair it is unsure about; precision and recall on the labelled corpus are both 100%. The linked source is authoritative.
  • “Independent source groups” is an estimate from publisher, wire credit, near-identical headlines and shared verbatim passages. When it cannot tell, it says so and does not count the reports as independent.
  • Tamil is handled by a conservative suffix normaliser plus place / concept lexicons — not a full Tamil NLP system. Tamil ↔ English matches require a shared district, a compatible date and a shared entity or action; a shared “Tamil Nadu” alone never merges. The original Tamil text is always kept.
  • Metadata differences between reports are not claims of contradiction; only a genuine semantic conflict is marked “disputed”.
  • Feeds go down. When they do, IFFA keeps the last known good snapshot, marks it stale, and never shows “LIVE”.
  • IFFA is not an emergency service. For any emergency, follow the issuing authority’s own instructions.

Refresh & data status

The live feed is regenerated on a schedule by a GitHub Actions workflow that runs the ingestion, validates the output, rebuilds the static site and redeploys it. The last snapshot in this build was generated 23 Sept 2026, 19:12 IST with 34 of 41 feeds responding.

See the source directory for per-feed status, or the methodology demonstrations for synthetic worked examples of the comparison model.