Model card
| Model | AI4M malaria risk model |
| Name | AI4M Malaria Model 1.0 |
| Version | ai4m-malaria-1.0, as returned in model_version |
| Type | Gradient-boosted decision trees (LightGBM), regression |
| Predicts | The share of malaria rapid diagnostic tests expected to be positive, from 0 to 1 |
| Coverage | Nigeria: 37 states, 774 LGAs, monthly from January 2018 |
| Maintained by | The AI4M team, [email protected] |
Intended use
Section titled “Intended use”To estimate how common malaria infection is likely to be in each LGA and state of Nigeria in a given month, so that public health teams can compare areas, watch trends and plan ahead.
It is meant to be one input to planning and prioritisation, alongside routine surveillance and local knowledge.
Not intended for
Section titled “Not intended for”- Confirming an outbreak. The model shows where conditions favour malaria. It does not see reported cases, so confirming that an outbreak is under way is a job for surveillance.
- Clinical or individual decisions. It says nothing about any one person or facility.
- Deciding on its own where resources go. A score should not be the sole reason to give or withhold support.
- Use outside Nigeria, or for periods shorter than a month.
What goes in
Section titled “What goes in”| Input | Source |
|---|---|
| Rainfall this month, last month, and the average of the last three | CHIRPS |
| Rainfall typical for the calendar month, and the difference from it | CHIRPS |
| Population density | WorldPop |
| Urban or rural | Derived from population density |
| Malaria positivity measured nearby in the previous national survey, within about 25, 50 and 100 km | The DHS Program surveys |
| Childhood anaemia and haemoglobin, net use, household wealth and electricity measured nearby in the latest of those surveys | The DHS Program surveys |
See Data sources.
What it learned from
Section titled “What it learned from”Malaria rapid diagnostic test results from four national household surveys: the Nigeria Malaria Indicator Surveys of 2010, 2015 and 2021, and the 2018 Demographic and Health Survey. In these surveys, children aged 6 to 59 months are tested in sampled clusters of households.
After leaving out clusters with fewer than five tests or missing inputs, the model was trained on 2,235 survey clusters across all 37 states.
It was not trained on health-facility case reports.
How it was tested
Section titled “How it was tested”- Forward in time. Each of the 2018 and 2021 surveys was predicted by a model trained only on the surveys before it.
- With no input that contains the answer. No input is derived from the survey being predicted.
- Against simple methods. It was compared with four simpler methods, including using what the previous survey measured nearby.
- On states left out. Whole states were withheld to see how well the model carries over to places it has not seen.
- With labels shuffled and with ten random seeds, to check the result is real and stable.
- For calibration. A correction was learned from the 2018 test and checked on the 2021 test.
- Across a longer gap. The 2021 survey was predicted using survey results no later than 2015, to match how far the service now runs past its latest survey.
- Against monthly reported cases. In the three states where facility data is public, monthly scores were compared with reported malaria cases for 2018 to 2024.
- For its error by area. Predictions were compared with survey results for whole LGAs and states.
How accurate it is
Section titled “How accurate it is”In plain terms
Section titled “In plain terms”On a scale of 0 to 1, comparing estimates with what later surveys measured:
| Area | 8 in 10 estimates were within | Average error |
|---|---|---|
| A state | 0.13 of the surveyed value | 0.09 |
| An LGA | 0.24 of the surveyed value | 0.15 |
These are the ranges returned as score_low and score_high.
Accuracy is modest. The model is useful for ranking areas and for giving every LGA a monthly figure. It is not precise for any one area.
Validation figures
Section titled “Validation figures”R² is the share of the variation between places that the model explains: 0 is no better than guessing the average everywhere, and 1 is perfect.
| Test | Result | Best simple method |
|---|---|---|
| Trained on 2010 and 2015, tested on 2018, individual survey locations | R² 0.24, average error 0.20 | R² 0.12 |
| Trained on 2010, 2015 and 2018, tested on 2021, individual survey locations | R² 0.29, average error 0.17 | R² 0.09 |
| The same tests, for whole states | R² 0.47 and 0.45, average error 0.08 | R² 0.39 and 0.35 |
| How well it ranks states from highest to lowest risk (1 is perfect) | 0.73 and 0.55 | 0.70 and 0.53 |
| States withheld, compared with a random split | R² 0.23 against 0.33 | |
| Telling high-risk locations from the rest (1 is perfect, 0.5 is chance) | 0.76 and 0.77 | 0.70 and 0.68 |
| Infected children found in the top-ranked fifth of locations | 29% and 30% | 27% and 29% |
| A six-year gap: 2021 predicted from 2015 survey results | R² 0.26, states 0.41 | R² −0.06 |
Between individual locations the model does two to three times as well as the best simple method, and for whole states it does somewhat better than using what the previous survey measured nearby. The 2021 figures include the calibration correction; the 2018 figures cannot, because the correction is learned from that test.
Calibration
Section titled “Calibration”Scores are corrected so that they match what surveys go on to measure. On the 2021 test, after a correction learned from the 2018 test, the calibration slope was 1.00 (1 is ideal) and the average error in level was under 0.01. Across ten groups from lowest to highest predicted risk, predicted and measured positivity agreed to within about 0.03.
Month-to-month movement
Section titled “Month-to-month movement”In Kwara, Nasarawa and Zamfara, the seasonal pattern of the monthly state score agreed with reported malaria cases at 0.50, 0.26 and 0.53 (1 is perfect agreement), and reported cases followed the score by two to three months. This is weak: the model learns differences between places far better than movement within a year. For the timing of the season, use the transmission season, which agreed with independent records at 0.88.
Read these with the limitations below. The figures measure how well the model carries over from earlier surveys to a later one. They do not measure how well it tracks change from month to month, because there is no data to test that.
Why the Malaria Atlas Project’s malaria maps are not inputs
Section titled “Why the Malaria Atlas Project’s malaria maps are not inputs”Those maps are built from the same surveys the model is tested against, so a map for a survey year already contains that survey. A model using them scores far higher on a test than a live service could. AI4M uses what earlier surveys measured instead.
What drives a score
Section titled “What drives a score”The inputs with the most influence are population density, the positivity measured within about 25 km in the previous survey, children’s haemoglobin measured nearby, and whether an area is urban or rural. Their effects run in the directions known from malaria epidemiology in Nigeria: denser, more urban areas have lower scores, and areas where the last survey found more infection have higher scores.
Known limitations
Section titled “Known limitations”- Outbreak prediction is delivered in part, through risk. Risk is estimated for every LGA and month, and indicates where outbreaks are more likely (it separated outbreak months from others at 0.83 across three states, where 1 is perfect), though not when one will begin. Case forecasts and outbreak probabilities work in three states; early warning of new outbreaks nationally needs monthly facility data, which AI4M does not yet have.
- Accuracy is modest: about a quarter of the variation between survey locations.
- Change over time is the least-tested part. Two tests years apart, none after 2021, surveys cover August to December only, and the model now works from survey results five years old.
- Local detail is limited. 82 of the 774 LGAs had no usable survey location in any year.
More in Limitations.
Ethical considerations
Section titled “Ethical considerations”A score is a statistical association at the level of an area. It is not a diagnosis and not a judgement on any community. Differences between neighbouring areas can be smaller than the model’s error, so rankings near each other should not be over-read. People affected by a decision informed by AI4M should have that decision reviewed by someone with local public health knowledge.
Updates
Section titled “Updates”The model’s inputs are refreshed about monthly as new rainfall data arrives. The model itself will be retrained when a new national survey is released. The model_version returned by Periods changes when it is.