The Monitoring Gap Your Clinical Staff Is Already Failing to Fill
The nurses at Morehouse School of Medicine were burning out. Two of them, managing follow-up calls for inflammatory bowel disease patients between clinic visits, hit a volume wall. The organisation's solution was to move patients into Zoom group sessions. Patients then sat on video calls discussing their bowel symptoms, pain levels, and medication side effects in front of each other. The privacy problem was obvious. The underlying economic problem was not: there was no model that could justify the cost of individual nurse calls at scale, and no technology that could substitute for them without creating new harms.
This is not an unusual situation. It is the default state of chronic disease monitoring in healthcare systems everywhere. A primary care physician managing thousands of patients can realistically engage with only dozens per day. The patients in the middle, not sick enough to justify a specialist appointment, not well enough to be ignored, exist in a monitoring gap that organisations fill with workarounds, burnout, and privacy-compromising compromises.
Researchers at IBM T.J. Watson Research Center, in collaboration with Cleveland Clinic Foundation and Morehouse School of Medicine, published a paper in July 2025 addressing this directly. Their system, Agent PULSE (Patient Understanding and Liaison Support Engine), is a telephonic AI agent that conducts medical surveys and continuous patient monitoring through natural conversation over standard phone lines. A pilot study with 33 IBD patients found that 70% expressed acceptance of AI-driven monitoring, with 37% actively preferring it over other modalities. The economic model the paper presents does not just argue that AI is cheaper: it identifies the precise severity zone where human intervention is economically unjustifiable and AI is the only viable option. That distinction matters more than the cost comparison.
Why the Severity Threshold Model Changes the Strategic Question
Most healthcare AI discussions frame the question as substitution: can AI replace a nurse, a physician, a call centre agent? The IBM paper frames it differently. It defines a tiered severity model where different care levels are economically appropriate at different points on the severity spectrum.
When a patient's condition exceeds a high severity threshold, specialist physician care is required. For moderate severity, nursing care is appropriate. For lower severity, untrained caregivers or family members typically provide support. Below the lowest threshold, for mild conditions that still warrant monitoring, the economic case for any trained human intervention collapses entirely. That is the zone Agent PULSE is designed for.
The insight is not that AI is cheaper than nurses. It is that below a defined severity threshold, the economics make continuous human monitoring structurally impossible, and without an AI alternative, patients in that zone receive nothing.
The paper formalises this with an efficiency ratio formula comparing human and AI total costs, where AI carries high fixed costs but near-zero marginal cost per additional patient. The authors present the formulas without actual dollar figures, which is an honest limitation they should get credit for acknowledging rather than obscuring. The economic argument is structurally sound even without the numbers, because the logic of digital marginal costs is well established. What the paper does not yet prove is the specific dollar threshold at which any given health system crosses from human-justified to AI-justified monitoring.
The Three Places the System Is Smarter Than It Looks
Agent PULSE runs over standard telephone lines, not a smartphone app or web portal. That design decision is not a technical constraint. It is a deliberate response to access economics.
Smartphone-based health apps require a device, an internet connection, sufficient digital literacy to navigate an interface, and enough visual or physical capability to operate a touchscreen. Standard telephone access is close to universal, including among elderly populations, low-income communities, rural patients with unreliable internet, people with visual or physical impairments, and patients with limited English proficiency. Voice is the only interface modality that does not impose a technology literacy requirement on top of a health literacy requirement.
The second design choice worth attention is the SOLOMON framework, IBM's proprietary multi-agent reasoning system sitting inside Agent PULSE. SOLOMON converts unstructured conversational transcripts into structured questionnaire responses. A patient saying "I've had a rough few days, my stomach's been awful and I can barely get off the couch" produces a usable MHBI severity score on a physician dashboard, not a voice recording that a nurse then has to review and manually code. The system handled speech-to-text transcription errors, too: when one patient said "loading" instead of "bloating," the LLM inferred the intended meaning from context and maintained conversational flow.
The third choice is what the system does when it finds something concerning. Patients can explicitly request a callback from a human provider mid-conversation. The system also detects concerning patterns autonomously and triggers escalation alerts. Autonomous routine assessment, human escalation for anything above the severity threshold: that division of labour is the operational core of the economic model.
What the Survey Completion Data Actually Reveals
The most operationally important finding in this paper is not the 70% acceptance rate. It is buried in the results section and largely absent from the conclusion.
Completion rates for questions about daily activities and immediate symptom impact reached 94.4%. Questions appearing later in the survey, including environmental triggers, treatment feedback, and research-related items, dropped to under 10%.
Patients are not treating AI check-ins with the same social obligation they would apply to a call with a human provider. Environmental distractions, one patient was driving through a toll booth mid-call, reduced engagement further. The paper's authors describe this honestly: "patients interacted with the AI system differently than they would with human providers," and they flag the trade-off between authenticity and completeness as warranting further investigation.
For any healthcare organisation considering deployment, this data point has direct operational consequences:
- The most clinically critical questions must be placed at the beginning of any survey instrument, not distributed across a long interaction
- Survey length is a clinical design constraint, not just a UX preference
- Environmental triggers and treatment barriers, often the most actionable information for care teams, will be systematically underrepresented if question sequencing follows standard questionnaire logic rather than AI engagement logic
- The 94.4% completion rate for symptom questions is genuinely encouraging; the sub-10% rate for later questions is a design problem that can be solved, but only if it is treated as one
The paper's authors frame this as future work. Organisations deploying voice health agents before that future work is done will be building dashboards on incomplete data and may not know which fields are systematically missing until they audit the completeness patterns.
The White-Coat Effect Works in Reverse Here
One finding the pilot surfaced is worth tracking carefully. The authors note that the absence of social pressure from a human provider may have allowed patients to express themselves more authentically, potentially revealing more accurate information about their conditions. They call this the inverse of the white-coat effect.
If the finding holds across larger populations and conditions, it has a specific clinical implication: AI-gathered symptom data may be more accurate for certain categories of sensitive or stigmatised information (mental health symptoms, substance use, treatment non-adherence, sexual health) than human-gathered data, not because the AI is better at asking questions, but because patients are less likely to perform wellness in front of it.
The 33-patient pilot cannot establish this. But it is the kind of hypothesis that, if confirmed at scale, would reposition AI voice monitoring from a cost-saving compromise to a clinically superior data collection method for specific question domains. That is a meaningfully different strategic proposition.
Where This Sits in the AI Health Landscape Right Now
The paper is a position paper with illustrative pilot data. It is not a randomised clinical trial. The 70% acceptance figure comes from 33 patients at a single site, all of whom had previously used Zoom group sessions, a modality with a known and significant privacy problem. When the comparison baseline is "discussing your bowel symptoms on a group video call with strangers," individual AI monitoring has a structural advantage that would not generalise to every deployment context.
The economic model presents formulas without actual cost estimates. The KV cache optimisation figures cited in the technical appendix (2.2 to 5x inference throughput improvements from CacheBlend, 3.2 to 3.7x latency reduction from CacheGen) come from external research papers, not from Agent PULSE itself. They are presented as aspirational targets, not measured results.
The IVR comparison was explicitly excluded by clinical partners. That exclusion is presented in the paper as evidence of IVR's market failure, which is a reasonable interpretation. It also means the study cannot demonstrate superiority over the class of systems it is most directly replacing.
What the paper does establish, within its scope, is a working architecture that could be deployed over standard telephone infrastructure, tested against validated clinical instruments, and accepted by a majority of patients in a chronically underserved population. That is not nothing. For health systems evaluating AI monitoring for chronic disease populations, it is a credible proof of concept with clear design parameters for the next study.
The Real Stakes Are Not Cost Savings. They Are Coverage.
Healthcare AI is typically sold on efficiency: fewer nurse hours, lower call centre costs, faster triage. The IBM paper's economic model supports that framing, and the authors use it. But the more durable argument is different.
For mild-severity chronic disease patients in underserved populations, the alternative to AI monitoring is not a nurse. It is nothing. The monitoring gap is not a cost inefficiency to be optimised. It is a structural absence. Morehouse School of Medicine's IBD patients were getting group Zoom calls because individual monitoring was economically impossible. Before that, two nurses were burning out trying to provide it. The AI agent is not cheaper than the nurse. It is the only option that exists at the coverage level these patients need.
As AI voice infrastructure matures, question ordering gets solved, EHR integration arrives, and longitudinal outcome data accumulates, the economic and clinical case will sharpen. Health systems and insurers waiting for that evidence before beginning pilot design will be 18 to 24 months behind organisations that are already building the operational infrastructure and learning which question sequences produce complete data.
The gap is not going away. The only question is what fills it.
Agents Applied is published weekly for C-suite executives, founders, and senior technology strategists navigating the operational realities of AI deployment.