How this briefing is made
This system reads Nepali news media daily, discards everything that is not about health, classifies what remains against Nepal’s own surveillance instruments, scores how much each report can be trusted, and groups reports of the same event. An analyst then decides what enters the briefing. Nothing publishes automatically.
Every figure below is read from the same configuration the pipeline runs on, so this page cannot drift from what the software actually does.
What classifies, and what writes
Classification uses TypeSafe’s Jev, a model that returns typed answers with calibrated probabilities rather than generated text. It is asked closed questions: choose one hazard class from a defined list, score severity against a written rubric, answer yes or no to a specific fact. It therefore cannot invent a category that does not exist, and when it is uncertain the probability distribution says so.
Because Jev generates no text, it also cannot write a quotation. Quotation spans are located in the source article first, and Jev only chooses among them. Each published excerpt is then checked to be a literal substring of the fetched page; anything that fails prints as UNVERIFIED.
A separate model translates the chosen Devanagari excerpt into English and renders one caption line from already-extracted fields. It is never asked to summarise across items or to write an executive paragraph, because a cross-item trend is the one claim in a briefing that no reader could check.
The credibility score
Credibility answers “how much should I trust this report”. It is kept strictly separate from severity, which answers “how serious is the event”. A report can be impeccably sourced and trivial, or alarming and unverifiable, and collapsing the two would hide exactly the cases that need an eye.
+ (25% × quantitative specificity)
+ (25% × outlet tier)
+ (15% × corroboration)
− penalties + bonuses, clamped to 0–100
The four components are always printed beside the score, in the interface and in the exported document. A score without its parts would be an assertion; with them it is an argument you can disagree with.
Source proximity
How close the central factual claim sits to an authoritative source.
| 0 | Unattributed, anonymous, or hearsay |
| 1 | Unnamed official, or attribution to another news outlet |
| 2 | Named health worker, named local official, or the outlet’s own reporter on the ground |
| 3 | Named senior clinician, named facility head, or named district health office |
| 4 | A government body or ministry speaking officially, or published surveillance data |
Quantitative specificity
Whether the report gives figures a reader could check. This is where evasion is caught: a story that declines to give a number where a number plainly belongs scores low here and is penalised again below.
| 0 | No figures at all. Only vague words such as many, several, or increasing |
| 1 | A vague magnitude, such as dozens or hundreds, with nothing checkable |
| 2 | Concrete figures, but no time window and no denominator |
| 3 | Concrete figures with a clear time window |
| 4 | Full figures with a time window, a denominator, and a breakdown |
Outlet tier
A static per-domain judgement about editorial standards, held in configuration rather than asked of a model, because a publication’s reputation is not a property of any single article. For a publisher absent from the list, the classifier estimates a tier from the writing and sourcing visible in the article alone, and the briefing marks that estimate as such.
| Tier | Portals | Who |
|---|---|---|
| 4 | 6 | BBC Nepali, Kantipur, Nagarik News, Onlinekhabar, Setopati, The Kathmandu Post |
| 3 | 12 | Annapurna Post, Baahrakhari, Gorkhapatra, Health TV Online, Himal Khabar, Naya Patrika, Nepali Times, Ratopati, Republica, Swasthya Khabar, The Himalayan Times, The Rising Nepal |
| 2 | 17 | Bizmandu, Clickmandu, Dainik Nepal, Deshsanchar, Drishti News, Karobar Daily, Kathmandu Pati, Khabarhub, Lokaantar, Nepal Khabar, Nepal Press, Nepalnews.com, News24 Nepal, Pahilopost, Rajdhani Daily, Thahakhabar, Ujyaalo Online |
| 1 | 3 | Image Khabar, Nepal Live, Reporters Nepal |
Corroboration
How many independent portals carry the same event. This saturates logarithmically at 12, because the difference between one portal and two is evidence, while the difference between eleven and twelve is a wire story being reprinted and says nothing further about whether the event occurred.
Penalties and bonuses
Applied after the weighted sum, each scaled by the probability the classifier assigned to that judgement rather than as an on-or-off switch.
| Signal | Effect |
|---|---|
| rumour | −30% |
| evasive | −18% |
| sensationalised | −12% |
| contradicts official | −12% |
| fluff promotional | −10% |
| ground reported | +6% |
| government response stated | +4% |
How signals are ranked
+ (35% × surveillance value/4)
+ 0.5 if any promotion rule fires
+ 0.08 if out of documented season
− up to 0.35 per demoting flag
× the multiplier for what kind of item it is
What kind of item it is
Every signal is classified on a second axis: not what it is about, but what kind of thing it is. Cases, clusters, deaths, service failures and exposures are events, which is what surveillance is for. A vaccination campaign, an inspection or a new guideline is a response. An explainer or a prevention feature is background. A ribbon-cutting or a political row is noise.
The axis exists because it was missing. A notice that Japanese encephalitis vaccination would begin in Jhapa led a briefing under a notifiable-disease dagger, and every part of that classification was correct: the hazard, the condition, the sourcing, the Gazette listing. Nothing had happened. The promotion rules keyed on the disease, and a disease is not an occurrence.
Two things follow. The promotion rules below now require an event, so a campaign against a notifiable disease is no longer promoted by rule. And the final priority is multiplied by the weight for its class, applied last so that a well-sourced conference cannot climb past a real death by accumulating bonuses.
| Class | Multiplier | What it covers |
|---|---|---|
| Event | ×1.00 | Outbreak, single case, service failure, exposure |
| Context | ×0.45 | Response, campaign, policy, appointment, procurement |
| Background | ×0.25 | Advisory, awareness, explainer, study, annual statistics |
| Noise | ×0.10 | Ceremonial coverage, political dispute, rewritten press release |
Surveillance value is asked of the model directly and separately from severity, because they are different questions. A single confirmed measles case in a country pursuing measles elimination is low severity and high value. A well-written feature on diabetes is the reverse.
Promotion overrides the score
Any of the following sends an item to the top of the briefing regardless of how it scored, with the reason printed beside it. The asymmetry is deliberate: burying a cholera case because its sourcing scored badly is a worse failure than showing one item too many.
All of them require the item to be an event. None fires on a response, an advisory or a ceremony, however serious the disease it names.
- Notifiable under Nepal’s 2024 Gazette list of 52 diseases
- Notifiable in all circumstances under IHR (2005) Annex 2
- Reportable under the EWARS standard operating procedure
- One or more human deaths reported
- A disease under a national elimination target
- Severity at the national-response level
Demotion keeps, never deletes
Political attacks on the health administration, promotional filler, and general health advice are ranked down rather than discarded. They stay in the file, in the appendix, so the filter can be audited. The same applies to anything the analyst excludes: it appears struck through rather than removed.
Sections
The top 15 items appear in full detail. The next 25 appear compressed to three lines. Everything else appears as one appendix line. Nothing is dropped.
The taxonomy
Two axes. One mutually exclusive hazard class, then a specific condition resolved within that class. The axes are separate because overlapping options split probability mass between them and depress confidence: a dengue story asked to choose between “outbreak” and “vector-borne” splits, because both are true.
Hazard classes follow WHO’s current scheme from the Global Public Health Intelligence Report 2022. Two classes are additions, marked below, because a Nepali national briefing must carry health-system failure and chronic environmental exposure and WHO’s media-monitoring scheme does not name them.
| Class | WHO category |
|---|---|
| Infectious | Infectious |
| Animal or zoonosis | Animal or zoonosis |
| Food safety | Food safety |
| Chemical | Chemical |
| Radiological or nuclear | Other / Radionuclear |
| Disaster or mass casualty | Disaster |
| Product | Product |
| Noncommunicable or injury | Other / Noncommunicable |
| Nutrition | Other / Nutritional deficiency |
| Health system | Addition. not a WHO GPHI category |
| Environmental | Addition. WHO folds this into Chemical and Other |
| Societal | Other / Societal |
| Undetermined | Other / Undetermined |
| Not health | n/a |
Conditions are grounded in Nepal’s own instruments rather than a generic disease list: 98 conditions in total, 52 from the Government Gazette notifiable list of 4 July 2024 under the Public Health Service Act 2075 s.49(1), 18 reportable diseases and 8 reportable syndromes from the EWARS standard operating procedure 2025, 4 notifiable in all circumstances under IHR Annex 2, and 10 under a national elimination target.
Each condition carries the Devanagari term Nepali journalism actually uses, which is often not the clinical term. Kala-azar appears as कालाजार in the press but भिसेरल लिस्मानियासिस clinically; diphtheria as भ्यागुते रोग; typhoid as म्यादी ज्वरो. Matching only the clinical form would miss the story. Of 98 conditions, 63 carry a term observed in live Nepali health journalism and 35 carry our own transliteration, which is the set most likely to need correction.
Geography
Signals resolve against Nepal’s 7 provinces, 77 districts and 753 local levels, to ward where an article specifies one. Two traps are handled explicitly. Nepali articles open with the city they were filed from, usually Kathmandu, which is usually not where the event happened, so dateline mentions are marked and the classifier is asked to reject them. And many district names are shared with their headquarters municipality, or with local levels in other districts, so an ambiguous name that nothing in the article disambiguates is reported as UNRESOLVED rather than guessed.
Deduplication
One wire story can appear on thirty portals within an hour. Articles are fingerprinted with a 64-bit SimHash over four-word shingles. A Hamming distance at or below 12 is treated as the same event outright. Between that and 20, the pair is put to the classifier as a single yes or no question, and a probability at or above 0.7 merges them. Every pairwise verdict is stored with its probability.
Sources
38 portals are enabled of 43 inventoried: 19 read through a feed, 16 through sitemaps, and 3 through a real browser because they do not serve plain HTTP clients.
Publishers are read at human pace. One request in flight per domain, spacing that is jittered rather than fixed, a long pause after a short run of requests, a per-run request ceiling, and conditional requests so an unchanged feed costs the publisher a 304 with no body. Any Crawl-delay a publisher declares always wins over our own floor. The crawler sends one honest user agent with a contact address and never rotates or disguises it.
Where a publisher serves only a rendered browser, a real browser is used, with the page’s own scripts executing so their advertising is served as it would be to any reader. That path is paced slower still and never used for archive collection.
What this cannot tell you
- Whether an event actually happened. Only that a named outlet reported it, and how well sourced that report appears.
- Anything a Nepali news outlet did not publish. Under-reported districts stay under-reported here.
- A trend or a baseline, until enough history exists to support one. No trend is shown that the archive cannot carry.
- A case count that means what a surveillance return means. A figure here is a journalist’s figure, carried verbatim with its source named.
Below 0.5 probability of being health-related, an article is not treated as a signal at all. Below 0.6 confidence on any classification, a signal is flagged for review before it is trusted.