Ailoitte LOGO
Case Study · AI Agent Development

Building an AI agent that remembers for a connected-care platform

Category
Healthcare · connected-care platform (50M+ members, anonymized)
What we built
A persistent-memory care coordination and follow-up agent
Core capability
Episodic + semantic memory · RAG grounding · preference learning
Delivery
AI Velocity Pod · ISO 27001 · SOC 2 Type II · HIPAA-ready
 persistent-memory care coordination and follow-up agent
62%
lower cost per interaction, from re-sending full history to retrieving only relevant context
14-day
context retained per member, up from a single session that reset on every call
<1%
ungrounded clinical claims, down from an unguarded baseline with no record grounding
31%
more follow-ups completed on time once coordinators stopped re-establishing context

About the project

Care Coordination Network

The client

A US connected-care platform that links payors, providers, and pharmacies for tens of millions of members, running high-volume care coordination and follow-up across chat, voice, and web.

The challenge

Their existing assistant started every session cold. Members re-explained their history, coordinators repeated context, and multi-day workflows lost the thread. That drove up cost, weakened personalization, and eroded trust in a setting where a wrong recollection has real consequences.

Multi-day Member Context
Multi-day Member Context

The approach

We treated memory as a first-class, governed subsystem, not a prompt trick. Four memory types work together behind a retrieval layer, so the agent recalls what matters, grounds answers in the member's own records, and improves with use while staying inside a HIPAA-ready boundary.

The impact

The outcome the client measured: more follow-ups completed on time, because coordinators picked up multi-day cases where they left off instead of rebuilding context. Members stopped re-explaining themselves, cost per interaction fell from full-history prompts to retrieval-only, and every clinical statement stayed traceable to a source record.

How we built it

The agent serves members, care coordinators, and clinicians, so the build had to keep one bounded, auditable memory system behind native and web surfaces, with retrieval doing the heavy lifting instead of a bloated context window.

Engagement · AI Velocity PodInterop · HL7 FHIR + MCP
Step 1

Memory model & boundary

  • Define what the agent may remember, for how long, and for whom
  • Partition memory per member with role-based access and consent scoping
  • Set the regulated boundary: what is a measurement, what is a suggestion
  • Open the audit and retention plan before writing a single record
Step 2

Retrieval & grounding

  • Episodic memory of prior sessions in an encrypted, per-member store
  • Semantic memory via RAG over FHIR records, embedded and filtered by authorization
  • Memory consolidation to summarize long journeys without losing key facts
  • Every clinical claim cites its source record; missing data triggers a handoff
Step 3

Evaluation & change control

  • Eval harness for recall accuracy, grounding, and alert precision
  • Guardrails against ungrounded claims, unsafe advice, and PHI leakage
  • Locked behavior where provable output is required
  • Predetermined change control plan for anything that learns, never silent updates

How the memory works

Memory is not one thing. Four layers do distinct jobs, and each is scoped, encrypted, and auditable so recall never becomes a compliance risk.

01Working memory
02Episodic memory
03Semantic memory (RAG)
04Preference & learning

Working memory holds the live turn. Episodic memory recalls prior sessions, so a returning member or coordinator never re-explains a multi-day case. Semantic memory retrieves from the member's own FHIR records with retrieval augmented generation, grounding every answer in real data instead of guessing. The preference layer tunes tone, channel, and reminder timing to the individual to cut fatigue.

Retrieval, not a giant context window, is what makes this affordable. Only the relevant slice of history reaches the model, so prompts stay short, latency stays low, and cost per interaction falls. Consolidation compresses long journeys into durable summaries without dropping the facts that matter.

Kept inside the boundary

Every model that touches a clinical fact, a recalled record, or an alert is treated as part of the regulated system: validated on representative data, bounded, documented, and governed by change control.

One member journey

A representative walkthrough of a single follow-up case, showing where each memory layer does its work. Names and details are anonymized and illustrative.

Day 1

First contact

  • Member messages about a newly prescribed medication and a side effect
  • Agent pulls the active care plan and prescription from FHIR (semantic memory)
  • Logs the concern, the guidance given, and an open follow-up task (episodic memory)
  • A low-confidence symptom question is handed to a human coordinator
Day 3

Returning contact

  • Member returns; the agent opens on the Day 1 concern with no re-explaining
  • It respects the member's preferred channel and quiet hours (preference layer)
  • Records that the side effect has eased and updates the follow-up task
  • Surfaces the pending lab the care plan still requires
Day 9

Coordinator handoff

  • A coordinator picks up the case and reads a consolidated summary, not a raw transcript
  • Consolidation kept the medication and lab history, dropped the small talk
  • Every clinical line links back to its source record for review
  • The loop closes: follow-up completed, task cleared, memory retained under policy
Background graphic
The human stakes

An agent that forgets makes a returning patient a stranger every time. Memory, done safely, is what turns a chatbot into a coordinator someone can trust.

Tech stack

Memory & retrieval
Vector storeEmbeddingsRAG pipelineEpisodic storeMemory consolidation
Orchestration
Multi-agent coordinationModel Context Protocol (MCP)Agent-to-agent (A2A)
Interoperability
HL7 FHIRSecure APIsEvent-driven triggersEHR & CRM connectors
Models
Model & framework agnosticLocked algorithms where provableOn-prem / private cloud option
Cloud · HIPAA-eligible
AWS, Azure, or GCP
Security & governance
Per-member partitioningEnd-to-end encryptionRole-based accessAudit loggingGuardrails & eval harness

Governed by design

Security and accountability are built into the memory system, not bolted on. The agent is aligned to the standards a regulated healthcare deployment requires.

HIPAA-readyMinimum-necessary PHI21 CFR Part 11 audit trailPredetermined change controlNIST AI RMFISO 27001ISO 9001SOC 2 Type IIEU AI Act aligned

Retention

Memory is kept only as long as the care relationship and policy allow. Each memory type carries its own retention window, agreed with the client, and expires automatically rather than accumulating by default.

Deletion & right to be forgotten

A deletion request purges a member's episodic and preference memory and any derived summaries, not just the underlying record. The removal is logged and verifiable, so nothing lingers in a cache or index.

Consent revocation

Revoking consent propagates to the memory layer immediately: retrieval for that member is cut off, and affected memory is quarantined or purged per the agreed policy, so stale context never keeps informing the agent.

Data residency

Memory stores stay inside the client's required region and cloud boundary, with an on-prem or private-cloud option where residency or air-gapping is mandated.

Our Legacy

Proof of Scale

Before we built our AI factory, we architected the platforms for some of the fastest-growing enterprises...

Assurecare screenshot
Assurecare logo
Assurecare

Connected Care Platform for 53M+ Members

Ailoitte helped power AssureCare’s patient-centered platform that connects payors, providers, and pharmacies to improve access, reduce cost, and strengthen care coordination.

Read Case Study
Dr Morepen screenshot
Dr Morepen logo
Dr Morepen

1M+ Customers, 40+ Years of Trust

We helped Dr. Morepen bring trusted preventive healthcare into a seamless mobile experience with reorders, subscriptions, and health tools.

Read Case Study
Utsah screenshot
Utsah logo
Utsah

Utsah Smart Ring

Your companion in a journey towards complete well-being, blending ancient Ayurvedic wisdom with advanced technology.

Read Case Study
iPatientCare screenshot
iPatientCare logo
IPatientCare

iPatientCare: an EHR built by physicians, for physicians

A patient-centered EHR software suite for primary care and specialty providers, built to enhance value-based care with a user-friendly experience.

Read Case Study

Frequently asked questions

What is a persistent-memory AI agent?
An agent that stores and retrieves context across sessions instead of starting cold each time. It combines working memory for the live turn, episodic memory of prior interactions, semantic memory retrieved from a knowledge base, and learned preferences, so it recalls history, grounds answers in real records, and adapts over time. It is a core pattern in our AI agent development services.
How is a memory agent different from a chatbot?
A chatbot responds turn by turn and forgets once the session ends. A memory agent plans across steps, calls tools, retrieves the right prior context, and acts with limited supervision. If your need is closer to scripted support, our conversational AI and chatbot development may fit; if it needs autonomy and recall, see AI agent development.
How is agent memory kept HIPAA-ready?
Memory is partitioned per member, encrypted at rest and in transit, access-controlled by role, and fully audit-logged. PHI is scoped to the minimum necessary, retrieval is filtered by consent and authorization, and every memory write and read is traceable for review. More on how we handle regulated data is on our security and compliance page.
How does the agent connect to our EHR and patient records?
Through HL7 FHIR, secure APIs, and event-driven triggers into your EHR, CRM, and data platforms, with retrieval filtered by authorization. This is standard in our healthcare software development work, including clinical documentation and EMR automation.
Why does memory reduce running cost?
Without memory, every session re-sends the full history into the model context, inflating tokens and latency. A retrieval layer sends only the relevant prior context, so prompts are shorter, responses are faster, and cost per interaction drops. We scope and cost these builds through fixed-price AI Velocity Pods.
How do you prevent it hallucinating patient facts?
Clinical answers are grounded with retrieval augmented generation over your own records, constrained by guardrails, and checked by an evaluation harness. The agent cites the source record for clinical facts and defers to a human when confidence is low or the record is missing. The grounding and model work sits within our generative AI and AI/ML development practice.
Should we build a custom agent or use an off-the-shelf one?
Off-the-shelf agents rarely survive contact with a regulated enterprise: they forget between sessions and are not governed for PHI. A custom agent is scoped to your data, systems, and compliance obligations, and you own the logic and roadmap. Our AI consulting team can map where a custom agent creates the most value.
How is a learning agent kept safe to update?
Anything that improves over time runs through a predetermined change control plan, never a silent update. Behavior is locked where provable output is required, and changes are validated and documented before release, aligned to the NIST AI Risk Management Framework. This governance is built into every agent we deliver.
How long does it take to build and deploy?
We move from discovery to a working proof of concept in weeks using a focused, cross-functional AI Velocity Pod with fixed-price, outcome-based scoping. A proof of concept validates value before a full build, so timelines stay predictable.
Can you extend our existing team instead of a full build?
Yes. Engage us for the full build or a single phase where you need specialist support, or hire dedicated AI developers to extend your team.
How do we get started?
Tell us the workflow you want to automate and the records it needs to reason over. We respond within twelve hours, and every conversation is protected by an NDA. Book a scoping call to map the memory model, grounding, and governance.

Want an agent that remembers, safely?

Tell us the workflow and the records it needs to reason over. We will map the memory model, the grounding, and the governance, and return a scoped plan.

Talk to an AI engineer

Recognized Leaders

logo

Top Innovative AI Companies 2025

TOI

Most Trusted IT Service provider 2024

International Business Times

The Best Software Development Company 2025

HT

Top 10 CEOs Share Their Vision for Success

logo

ISO 27001:2013 Information Security

AP NEWS

Enterprises scale teams faster

BS

Smarter Enterprises with Custom AI

logo

ISO 9001:2015 Quality Management