Skip to content

Model card

Model AI4M malaria risk model
Name AI4M Malaria Model 1.0
Version ai4m-malaria-1.0, as returned in model_version
Type Gradient-boosted decision trees (LightGBM), regression
Predicts The share of malaria rapid diagnostic tests expected to be positive, from 0 to 1
Coverage Nigeria: 37 states, 774 LGAs, monthly from January 2018
Maintained by The AI4M team, [email protected]

To estimate how common malaria infection is likely to be in each LGA and state of Nigeria in a given month, so that public health teams can compare areas, watch trends and plan ahead.

It is meant to be one input to planning and prioritisation, alongside routine surveillance and local knowledge.

  • Confirming an outbreak. The model shows where conditions favour malaria. It does not see reported cases, so confirming that an outbreak is under way is a job for surveillance.
  • Clinical or individual decisions. It says nothing about any one person or facility.
  • Deciding on its own where resources go. A score should not be the sole reason to give or withhold support.
  • Use outside Nigeria, or for periods shorter than a month.
Input Source
Rainfall this month, last month, and the average of the last three CHIRPS
Rainfall typical for the calendar month, and the difference from it CHIRPS
Population density WorldPop
Urban or rural Derived from population density
Malaria positivity measured nearby in the previous national survey, within about 25, 50 and 100 km The DHS Program surveys
Childhood anaemia and haemoglobin, net use, household wealth and electricity measured nearby in the latest of those surveys The DHS Program surveys

See Data sources.

Malaria rapid diagnostic test results from four national household surveys: the Nigeria Malaria Indicator Surveys of 2010, 2015 and 2021, and the 2018 Demographic and Health Survey. In these surveys, children aged 6 to 59 months are tested in sampled clusters of households.

After leaving out clusters with fewer than five tests or missing inputs, the model was trained on 2,235 survey clusters across all 37 states.

It was not trained on health-facility case reports.

  • Forward in time. Each of the 2018 and 2021 surveys was predicted by a model trained only on the surveys before it.
  • With no input that contains the answer. No input is derived from the survey being predicted.
  • Against simple methods. It was compared with four simpler methods, including using what the previous survey measured nearby.
  • On states left out. Whole states were withheld to see how well the model carries over to places it has not seen.
  • With labels shuffled and with ten random seeds, to check the result is real and stable.
  • For calibration. A correction was learned from the 2018 test and checked on the 2021 test.
  • Across a longer gap. The 2021 survey was predicted using survey results no later than 2015, to match how far the service now runs past its latest survey.
  • Against monthly reported cases. In the three states where facility data is public, monthly scores were compared with reported malaria cases for 2018 to 2024.
  • For its error by area. Predictions were compared with survey results for whole LGAs and states.

On a scale of 0 to 1, comparing estimates with what later surveys measured:

Area 8 in 10 estimates were within Average error
A state 0.13 of the surveyed value 0.09
An LGA 0.24 of the surveyed value 0.15

These are the ranges returned as score_low and score_high.

Accuracy is modest. The model is useful for ranking areas and for giving every LGA a monthly figure. It is not precise for any one area.

R² is the share of the variation between places that the model explains: 0 is no better than guessing the average everywhere, and 1 is perfect.

Test Result Best simple method
Trained on 2010 and 2015, tested on 2018, individual survey locations R² 0.24, average error 0.20 R² 0.12
Trained on 2010, 2015 and 2018, tested on 2021, individual survey locations R² 0.29, average error 0.17 R² 0.09
The same tests, for whole states R² 0.47 and 0.45, average error 0.08 R² 0.39 and 0.35
How well it ranks states from highest to lowest risk (1 is perfect) 0.73 and 0.55 0.70 and 0.53
States withheld, compared with a random split R² 0.23 against 0.33
Telling high-risk locations from the rest (1 is perfect, 0.5 is chance) 0.76 and 0.77 0.70 and 0.68
Infected children found in the top-ranked fifth of locations 29% and 30% 27% and 29%
A six-year gap: 2021 predicted from 2015 survey results R² 0.26, states 0.41 R² −0.06

Between individual locations the model does two to three times as well as the best simple method, and for whole states it does somewhat better than using what the previous survey measured nearby. The 2021 figures include the calibration correction; the 2018 figures cannot, because the correction is learned from that test.

Scores are corrected so that they match what surveys go on to measure. On the 2021 test, after a correction learned from the 2018 test, the calibration slope was 1.00 (1 is ideal) and the average error in level was under 0.01. Across ten groups from lowest to highest predicted risk, predicted and measured positivity agreed to within about 0.03.

In Kwara, Nasarawa and Zamfara, the seasonal pattern of the monthly state score agreed with reported malaria cases at 0.50, 0.26 and 0.53 (1 is perfect agreement), and reported cases followed the score by two to three months. This is weak: the model learns differences between places far better than movement within a year. For the timing of the season, use the transmission season, which agreed with independent records at 0.88.

Read these with the limitations below. The figures measure how well the model carries over from earlier surveys to a later one. They do not measure how well it tracks change from month to month, because there is no data to test that.

Why the Malaria Atlas Project’s malaria maps are not inputs

Section titled “Why the Malaria Atlas Project’s malaria maps are not inputs”

Those maps are built from the same surveys the model is tested against, so a map for a survey year already contains that survey. A model using them scores far higher on a test than a live service could. AI4M uses what earlier surveys measured instead.

The inputs with the most influence are population density, the positivity measured within about 25 km in the previous survey, children’s haemoglobin measured nearby, and whether an area is urban or rural. Their effects run in the directions known from malaria epidemiology in Nigeria: denser, more urban areas have lower scores, and areas where the last survey found more infection have higher scores.

  • Outbreak prediction is delivered in part, through risk. Risk is estimated for every LGA and month, and indicates where outbreaks are more likely (it separated outbreak months from others at 0.83 across three states, where 1 is perfect), though not when one will begin. Case forecasts and outbreak probabilities work in three states; early warning of new outbreaks nationally needs monthly facility data, which AI4M does not yet have.
  • Accuracy is modest: about a quarter of the variation between survey locations.
  • Change over time is the least-tested part. Two tests years apart, none after 2021, surveys cover August to December only, and the model now works from survey results five years old.
  • Local detail is limited. 82 of the 774 LGAs had no usable survey location in any year.

More in Limitations.

A score is a statistical association at the level of an area. It is not a diagnosis and not a judgement on any community. Differences between neighbouring areas can be smaller than the model’s error, so rankings near each other should not be over-read. People affected by a decision informed by AI4M should have that decision reviewed by someone with local public health knowledge.

The model’s inputs are refreshed about monthly as new rainfall data arrives. The model itself will be retrained when a new national survey is released. The model_version returned by Periods changes when it is.