A Think Piece for Public Health Stakeholders
Where Does AI Improve a Real Health Outcome1
Health systems across the Global South are not short of challenges and urgency. With community health workers (CHWs) and primary health centres (PHCs) stretched across rising non-communicable disease (NCD) burdens, climate-driven shocks, and population-scale demand, the data that could help: patient records, call logs, supply registers, weather feeds, sits fragmented across programs, geographies, entities and devices. Into this gap, AI is arriving fast, and increasingly with state backing. India’s Ayushman Bharat Digital Mission (ABDM), a critical Digital Public Good (DPG) infrastructure for identity, is crucial to bring AI and data together for a clearer view of who we are talking about, when and where. ABDM has linked over 859 million digital health IDs to more than 878 million health records by early 2026, and the government used the India AI Impact Summit in February 2026 to launch SAHI — a national strategy for AI in healthcare — and BODH — a benchmarking platform that lets AI models be validated against real-world health data before they reach the frontline. [1]
The Health Secretary’s framing at that summit is the one we find most useful: digital systems capture and transmit information; AI’s job is to support the providers and patients to make right decisions and choices. [1]
In this piece, we deliberately avoid the two most common failure modes in AI-for-health writing. The first is techno-optimism: cataloguing impressive pilots without asking whether they survived contact with a real PHC’s staffing and connectivity constraints. The second is reflexive caution: treating every AI use case as equally risky, when in fact, triage support, supply forecasting, and diagnostic image reading carry very different risk profiles and should be governed differently.
Our position is focused, and takes a view for Public Health Stakeholders: AI in health is worth adopting only where it demonstrably improves health outcomes (and processes leading to it) and the Providers (CHW, a PHC physician, a city health officer) — remains in the loop to decide. We call this: detect earlier, decide faster, deploy smarter. The test for any AI investment is whether it produces a measurable improvement in health outcomes, not whether it is technically impressive. As implementers and advisors who evaluate, adapt, and govern AI tools against exactly this test, rather than as generic technology adopters, this is the lens Swasti (https://swasti.org) and CMS (https://www.cms.org.in/) bring to their own work in this space. Their experience running and advising on the deployments below shapes the view that follows.
Beyond The Hypothetical: Four Categories Where AI Is Already Changing Health Outcomes
1 This piece is about Applied AI - practical AI tools that are already integrated into frontline worker and health-system workflows — rather than Frontier AI, which covers cutting-edge research and the development of new models. It therefore does not cover Frontier AI topics such as precision medicine, drug discovery, or similar areas, which sit more in the realm of pharmaceutical companies and large technology corporations.
The samples we have selected below have moved beyond lab demonstrations and pilot stages and are those that are deployed, evaluated, and at meaningful scale.
1. Predictive Risk Targeting: Using AI to Decide Who Needs a Human Touch First
ARMMAN, an Indian maternal-health NGO, runs mMitra – a voice-message programme reaching over 2.3 million women, but its call centre could only follow up with roughly 1,000 women a week against a far larger pool of mothers at risk of disengaging. [2]
Working with Google Research India, IIT Madras, and later Harvard, ARMMAN deployed a restless multi-armed bandit (RMAB) model: a reinforcement-learning approach built specifically for the problem of “who should we call this week given limited calling capacity”, to rank beneficiaries by dropout risk and likely benefit from a call. Field-tested results showed up to a 32% reduction in engagement drop-off among high-risk women, and a separate analysis attributed roughly a 30% improvement in retention to the prioritisation model. [3,4]
The deployment, now five years old and still running under what the research team calls the SAHELI architecture, is one of the most rigorously evaluated AI-for-health systems in the world precisely because it was built around a real operational bottleneck — limited CHW and call-centre hours — rather than around the technology itself. [4]
Why this matters: the constraint in most low-resource health systems is never data: it is human follow-up capacity. The highest-leverage AI use case in this category is not prediction for its own sake, but prediction that tells a small, overstretched human team exactly who to call, visit, or flag first. This pattern (rank and route, leave the decision and the conversation human) is replicable well beyond maternal health, into TB follow-up, NCD management, and post-discharge care.
2. Diagnostic Augmentation: Using AI to Multiply Scarce Specialist Capacity and ReduceRisks and Dropouts
India captures roughly 80 million chest X-rays a year for TB screening but does not have enough radiologists to read them quickly; manual reads also carry an error rate -Qure.ai’s own published material puts at 25–30%. [5] Their qXR tool, trained on over a million chest X-rays, reads a film in under 30 seconds and was endorsed by WHO in 2021 as usable without a human reader present. [6] In a PATH-led active case-finding programme in Nagpur, the presence of AI alongside radiologists increased overall TB case yield by 15.8%, because the model flagged presumptive cases that radiologists had not. [7] A 2025 health technology assessment commissioned by India’s Department of Health Research found qXR not merely effective but cost-saving relative to existing TB diagnostic pathways. [8] The tool is now deployed at over 2,600 sites in 67+ countries and reads an estimated 5–10 million X-rays a year, making it arguably the most scaled autonomous clinical AI use case in the world.[9]
Why this matters: this category works wherever a large volume of a single, structured signal (an X-ray, a lab result, a retinal scan) already exists, but the bottleneck is expert interpretation time rather than the underlying data. India’s own MadhuNetrAI tool for diabetic retinopathy screening follows the same logic, and is now part of the country’s national AI showcase [1] — a strong signal that this category is the most government-validated of the four.
This is also where Swasti's advisory work with governments becomes directly useful. As the number of AI and MedTech solutions grow, tools like qXR now sit inside a crowded vendor market while governments need independent capacity to evaluate which AI tool fits which problem, at what cost, and with what governance attached, rather than accepting vendor claims at face value: a gap BODH itself was built to help close.
Swasti applied this approach through its end-to-end advisory support to the Department of Health, Government of Andhra Pradesh, in convening the state’s first-of-its-kind MedTech Innovation Challenge. The challenge brought together 18 shortlisted organizations for a rigorous, evidence-based evaluation, with a high-level jury comprising healthcare leaders, clinical experts, senior government officials, and hospital administrators. The innovations were assessed not on the quality of their presentations, but on how well they performed in real-world settings, cost of ownership, feedback of users, the results they achieved, and their readiness to be deployed at scale. The challenge demonstrated the Department’s commitment to identifying solutions that can deliver measurable impact within public health systems. It also highlighted a larger need: governments need the systems and capabilities to evaluate, procure, govern, integrate, and monitor AI and MedTech solutions. This can help ensure that promising innovations do not remain isolated pilots, but are taken to scale where they can benefit larger populations.
3. Climate and Outbreak Early Warning — Using AI to Shorten the Lead Time on a Known Threat
Ahmedabad’s Heat Action Plan (HAP), launched in 2013 after a 2010 heatwave caused over 1,300 excess deaths, pairs a 7-day probabilistic weather forecast with graded alert protocols across hospitals, CHWs, and city agencies. [10] A peer-reviewed evaluation found that extreme-heat warnings issued under the plan were associated with measurably lower summertime all-cause mortality, with the largest reductions occurring at the highest temperatures — exactly when the system is needed most. [11]
The HAP has since been replicated in at least eight other Indian cities, though a 2025 review in the Indian Journal of Medical Research found significant unevenness: most of the newer plans copy Ahmedabad’s alert thresholds without adapting them to local mortality and morbidity data, and few link the warning system to long-term interventions like cool roofing or green space. [12]
Separately, machine-learning models — particularly LSTM-based deep learning — have shown strong performance forecasting dengue prevalence from climate variables such as humidity, temperature, and rainfall, well ahead of an outbreak becoming clinically visible. [13]
Swasti’s Precision Analytics and Action Towards Climate and Health (PATCH) initiative, run through its Precision Health Platform, is a useful illustration of where this category is heading next: a localised, action-linked early warning system that integrates environmental, health, and social-sector data into a single dashboard so that city and state governments can identify vulnerable areas, prioritise action, and shape risk communication in real time. PATCH is currently deployed with Tabaco City in the Philippines and with the state governments of Andhra Pradesh and Meghalaya in India, and is paired with technical support for evidence-based sensemaking rather than a dashboard alone. A related Swasti pilot in Ananthapuramu district, with the Government of Andhra Pradesh, paired this kind of predictive modelling — including a model to predict typhoid incidence from climate and health data at the block level — with frontline capacity building and a district Heat Action Plan vetted by India’s National Programme on Climate Change and Human Health. [17]
Why this matters: this category proves early warning saves lives, but also exposes its central weakness — the “first mover” version of a system rarely transfers cleanly to a new city without local recalibration to that city’s own mortality, morbidity, and vulnerability data. The opportunity is not building more forecasting models from scratch; it is the localisation, calibration, and last-mile linkage layer that turns a generic forecast into a ward-specific action protocol, and the shared infrastructure that lets multiple cities benefit from one underlying model. PATCH is an early, live example of exactly this pattern: a single underlying platform being calibrated and deployed across genuinely different geographies (a Philippine city, two Indian states) rather than each government commissioning its own model from zero.
4. Frontline Access and Decision Support: Using AI to Augment, Not Replace, the Community Health Workforce
Khushi Baby’s ASHABot, built with Microsoft Research India and launched in 2024, lets India’s roughly 1 million ASHA workers ask maternal- and child-health questions by voice note over WhatsApp in Hindi, English, or Hinglish, and receive evidence-based answers in seconds. [14]
This targets a structural problem CHWs have long reported: AI tools designed without their input often assume literacy, English fluency, or smartphone sophistication that frontline workers do not uniformly have. This finding was borne out in CHI-published research on CHW perceptions of AI-enabled mobile health tools in rural India, which found workers valued AI assistance but distrusted systems that could not explain their reasoning or that bypassed their judgment. [15]
Why this matters: this category is the access layer that determines whether any of the other three actually reach the last mile. The CHI research is a useful corrective against assuming CHWs will adopt any tool handed to them; trust, explainability, and a visible “the human still decides” design patterns are not governance nice-to-haves, they are adoption preconditions.
Swasti’s work in Andhra Pradesh illustrates how this principle can be applied beyond a chatbot. Swasti is implementing a digital pregnancy continuum and decision-support system across two districts in Andhra Pradesh, that brings together data from existing systems to create a longitudinal view of each pregnancy, from conception through delivery and postpartum care. The system helps frontline and district health teams identify high-risk pregnancies, track whether essential interventions have been delivered, and act on gaps in care. A high-risk pregnancy dashboard and analytics layer is developed, focused on conditions such as severe anaemia, PIH/pre-eclampsia, preterm birth and teenage pregnancy. As the data and system mature, AI and predictive analytics will be layered on to flag women who may be at higher risk, miss ANC visits, experience worsening anaemia, or face preterm birth or stillbirth — not to make decisions for health workers, but to help them know whom to prioritise, what requires attention, and where follow-up is most needed. This keeps AI in its appropriate role: augmenting the reach, prioritisation and decision-making capacity of the community health workforce while retaining human judgment at the centre.
Two Frameworks We Use to Decide Where AI Earns Its Place
Most AI-in-health writing stops short of a strategy at “here are some interesting use cases.” Once a category of impact is established, the harder and more useful question for a funder or a government is: given a specific proposed AI use case, should we invest, pilot cautiously, or wait? Swasti uses two simple frameworks to make that judgment, and offer them here as reusable tools.
1. The 4D Test: Detect, Decide, Deploy, Demonstrate
Any credible AI-in-health proposal should be able to answer four questions in sequence. If it cannot answer the first, it is not ready for the second.
| Stage | The question it answers | What “ready” looks like |
|---|---|---|
| Detect | Can the model reliably, in some cases better than human, identify the signal (risk, disease, threat) earlier or more accurately than current practice? | Validated against real outcomes, not just historical correlation; ideally peer-reviewed or government-benchmarked (e.g. via BODH). |
| Decide | Is there a specific human (a CHW, clinician, or official) positioned and equipped to act on the output? | A named decision-maker and workflow exist before the model is built, not bolted on after. |
| Deploy | Can the tool function inside real infrastructure (connectivity, language, device, staffing constraints)? | Field-tested in the actual operating environment, not just a lab or urban pilot site. |
| Demonstrate | Is there a measurable health outcome attributable to the tool, not just an output metric? | An independent impact evaluation design, not only an accuracy or adoption metric - but cost, data sharing fairness and security, field practicality and dependence of vendor. |
The discipline this enforces is sequential, not simultaneous: a tool that detects well but has no defined decision-maker (no CHW, clinician, or city officer positioned to act on the output) is not yet an investment-ready use case, it is a research project. Of the four categories above, predictive risk targeting and diagnostic augmentation are furthest along the 4D chain because their Detect and Decide steps are well-evidenced; climate and outbreak early warning is strong on Detect and Deploy but, per the Indian Journal of Medical Research review cited, frequently weak on Decide (local calibration) and Demonstrate (linked outcome data); frontline access tools are foundational to Deploy but rarely measured against Demonstrate on their own. [12]
2. The Readiness × Stakes Matrix
The second filter is a simple two-axis test we apply to any proposed AI use case: how mature is the underlying evidence and data (Readiness), and how severe are the consequences of a wrong or biased output (Stake)? The combination, not either axis alone, should determine the governance posture: a high-stakes, low-readiness use case (say, an unvalidated model recommending which patients to discharge) deserves far more caution than a low-stakes, high-readiness one (say, a well-validated model prioritising which low-risk mothers get a reminder call this week), even though both are “AI in health.”
| Stake | Low Readiness (new model / limited field evidence) | High Readiness (field-validated, multi-year evidence) |
|---|---|---|
| High Stakes (clinical, irreversible, or equity-sensitive decisions) | Don’t deploy. Research and validate first; sandbox or benchmark only (e.g. BODH-style testing). | Scale with governance. Human-in-the-loop, explainability, and audit trail are non-negotiable preconditions (e.g. qXR alongside a radiologist). |
| Low Stakes (recoverable, low-severity decisions) | Pilot with oversight. Reasonable to test in the field with close monitoring and a fast kill-switch. | Scale with monitoring. Lighter-touch governance is defensible; periodic outcome audits remain important (e.g. ARMMAN’s call-prioritisation model). |
Mapped against the four categories above: ARMMAN’s dropout-risk model sits in the Scale-with-monitoring quadrant: high readiness (five years of field evaluation) and moderate stakes (a missed call is recoverable). Qure.ai’s qXR sits closer to Scale-with-governance: high readiness and high stakes – which is exactly why WHO endorsement and regulatory clearance, not just technical performance, were preconditions for its scale. New heat or dengue models entering a city for the first time sit in Pilot-with-oversight: the underlying technique is proven elsewhere, but local readiness has to be earned before stakes can be safely raised. This is also why we are cautious about any proposal that starts in the bottom-right cell: a pattern that is unproven, high-stakes, and seeking to scale immediately, more than any other, is where AI-in-health pilots damage trust and set a field back.
Why Trust Architecture Is the Product And Not a Compliance Add-On
Every example above that successfully scaled did so because trust was not retrofitted but designed in. WHO’s 2021 guidance on AI ethics in health remains a key global reference and highlights six principles: protecting human autonomy, promoting human wellbeing and safety, ensuring transparency and explainability, fostering accountability, ensuring inclusiveness and equity, and promoting AI that is responsive and sustainable. [16]
ARMMAN’s model works because a human still makes the call; Qure.ai’s qXR works inside an existing clinical pathway with a radiologist still positioned as the authority in most deployments; Ahmedabad’s HAP works because it triggers a known, pre-agreed institutional response rather than an opaque automated action. Where the CHI research on CHW attitudes is most pointed is its finding that frontline workers’ trust in AI hinges, not on the tool’s technical sophistication, but on explainability and whether the tool visibly defers to their judgment. [15]
Swasti’s PATCH initiative shows that trust in AI-enabled public health systems is not built by technology alone. It is built by bringing the right people and institutions into the process from the beginning. In Andhra Pradesh, PATCH is being shaped around the state’s climate and health priorities, bringing together government health systems, weather and environmental data, and local decision-makers to turn climate signals into early warnings that can inform action. In Meghalaya, this approach goes a step further through a One Health model. PATCH brings together the Health and Veterinary Departments, the One Health Cell, the India Meteorological Department, space centres providing satellite-based ecological and land-use data, and academic and technical institutions. Each brings a different piece of the picture, helping the platform understand risks across human and animal health, climate, ecology and land use. Crucially, AI-generated predictions are not simply accepted at face value. They are explained, discussed with stakeholders and ground-validated, particularly in areas where the model identifies risks that may not be visible in routine reporting. PATCH creates multiple points where predictions can be questioned, tested against local realities and interpreted by people who understand the context. AI provides the signal; multi-stakeholder governance determines whether that signal can be trusted, what it means on the ground, and what happens next.
This is also where the government itself is now raising the bar. India’s SAHI framework and its BODH benchmarking platform exist precisely because procuring a validated, trustworthy AI tool has so far been harder than procuring an impressive-looking one. Any organisation seeking to operate in this space, whether as an implementer, an advisor, or a convenor, will increasingly be expected to demonstrate this kind of governance discipline, not as an afterthought, but as a precondition for funding and for government partnership. [1]
For donors and multilateral partners, the opportunity is not to fund another AI pilot when the field already has several well-evidenced ones. It is to fund the parts of the system that make those pilots transferable: localisation and calibration of existing models (heat, vector-disease, dropout-risk) to new geographies; the governance and consent infrastructure that lets community-level health data be used at all; and the independent advisory capacity that helps governments choose and validate AI tools rather than accept vendor claims at face value. For government stakeholders, the offer is a partner that has already done the work of separating what is proven — predictive risk targeting, diagnostic augmentation, climate and outbreak early warning, frontline access tools — from what is still speculative, and that will not deploy AI into frontline care without a measurable outcome case attached to it.
References
- India AI Impact Summit 2026 — Union Ministry of Health and Family Welfare panel on SAHI and BODH, February 2026 (digit.in; newkerala.com).
- ARMMAN — mMitra programme overview (armman.org).
- “ARMMAN scales its AI efforts to improve maternal and child health in India, with support from Google.org,” Business Standard / ANI–PRNewswire, February 2021.
- Verma, Dasgupta, Madhiwalla, Taneja, and Tambe, “Decisions and Deployment: The Five-Year SAHELI Project (2020–2025) on Restless Multi-Armed Bandits for Improving Maternal and Child Health,” arXiv preprint.
- Qure.ai, “Scaling up TB screening with AI: Deploying automated X-ray screening in remote regions,” Qure.ai blog.
- Qure.ai evidence pages — WHO endorsement of qXR for autonomous TB detection, 2021.
- Vijayan et al., “Implementing a chest X-ray artificial intelligence tool to enhance tuberculosis screening in India: Lessons learned,” PLOS Digital Health, 2023.
- Indian Institute of Public Health Gandhinagar (IIPHG), Health Technology Assessment of AI-assisted CXR interpretation for TB, cited via Qure.ai newsroom, 2025.
- “AI Chest X-ray: Revolutionizing Indian Radiology & TB Detection,” ocacademy.in, 2025.
- Natural Resources Defense Council, Ahmedabad Municipal Corporation, and Indian Institute of Public Health–Gandhinagar, Ahmedabad Heat Action Plan documentation.
- “Building Resilience to Climate Change: Pilot Evaluation of the Impact of India’s First Heat Action Plan on All-Cause Mortality,” PMC/NCBI.
- “Heat action plans in eight Indian cities: Knowledge gaps & opportunities for intersectoral heat governance,” Indian Journal of Medical Research, 2025.
- “Weather integrated multiple machine learning models for prediction of dengue prevalence in India,” International Journal of Biometeorology, 2022.
- “AI health chatbots: Reaching Rural Patients in India” (Khushi Baby’s ASHABot), The Borgen Project, 2025.
- Okolo, Kamath, Dell, and Vashistha, “‘It cannot do all of my work’: Community Health Worker Perceptions of AI-Enabled Mobile Health Applications in Rural India,” CHI 2021.
- World Health Organization, “Ethics and Governance of Artificial Intelligence for Health: WHO Guidance,” 2021.
- Swasti, “Climate x Health” — Precision Action Towards Climate and Health (PATCH) and Integrated Heat Program in Ananthapuramu (swasti.org/climate-x-health).