Healthcare

How Techparser Built AI Doctor, a Multimodal Medical Assistant

Techparser built AI Doctor, a multimodal medical assistant demo: voice and photo in, a streamed urgency badge out, with nothing written to disk.

Client
Techparser (live demo)
Industry
Healthcare
Timeline
Live demo
Team
1 senior full-stack engineer

Want results like these?




Results

3 steps
Speech-to-text, vision model, then text-to-speech
2 providers
Swappable behind one toggle per voice step
3 levels
Low, Moderate or High urgency badge

The problem

A person describes a health concern by voice, optionally attaches a photo, and gets back a structured assessment: an urgency level, possible causes, self-care steps and when to seek care, read aloud if they want. The conversation is multi-turn, so the photo and earlier answers stay in context for follow-up questions. It runs on serverless hosting, so nothing durable can live on the filesystem and concurrent users must never share state. Three model calls are chained, speech-to-text, a vision model and text-to-speech, yet it has to feel responsive, and two providers per voice step had to be swappable behind one switch.

The first version showed why this is hard. It passed audio and images through files on disk under shared names, which was broken on serverless and a way for one person's photo or recording to surface in someone else's session. Speech models also invent filler phrases on silent audio, which would turn into made-up symptoms. Both failures could put the wrong information in front of someone asking about their health.

What Techparser built

  • Audio arrives as in-memory form data and the image travels inside the chat message itself, so nothing is written to the server; that made it safe on serverless and impossible for concurrent users to collide.
  • The assessment streams token by token from the vision model, with proxy buffering switched off so the first words appear immediately, and a provider error mid-answer is appended as a short note rather than dropping the connection.
  • The model opens every reply with its urgency level, and the client turns that first line into a Low, Moderate or High badge while the rest renders as it arrives, hiding the half-typed line until it resolves.
  • Speech-to-text and text-to-speech each have two providers behind one toggle, and switching re-voices the latest reply, so cost and quality can be compared on exactly the same answer.
  • Transcripts are normalised and checked against the filler phrases speech models produce on silence; a match is treated as no speech rather than a symptom, and transcription runs at temperature zero.
  • The app's code is only sent to a visitor holding a valid signed session; everyone else gets the sign-in screen alone, and every model endpoint checks the session again on its own.
  • Sessions are HMAC-signed with an expiry, credentials are compared in constant time, and five failed attempts lock an address out for three hours; it fails closed if credentials are not configured.

Decisions that mattered

Stateless by design, nothing on disk

The first version passed audio and images through files on disk under shared names, which broke on serverless and let one person's photo surface in another's session. The rebuild keeps everything in memory: audio arrives as in-memory form data, and the image travels inside the chat message itself, so nothing is written to the server. That made the app safe on serverless and impossible for concurrent users to collide. The trade-off is that a conversation does not survive a reload, which is acceptable for a stateless demo and removes an entire class of cross-user data leaks.

Structure without losing the stream

A medical assessment needs structure, an urgency level and fixed sections, but structure usually means waiting for the whole response before parsing it. Instead the model is told to open every reply with its urgency level, followed by fixed sections. The client turns that first line into a Low, Moderate or High badge while the rest of the answer renders token by token, and the half-typed first line is hidden until it resolves. The decision was to encode structure in the output order rather than a separate parse step, so the reader gets an instant urgency signal and a streamed answer at the same time.

Not turning silence into symptoms

Speech models invent filler phrases on silent or near-silent audio, and in a health tool a hallucinated phrase becomes a made-up symptom. So transcripts are normalised and checked against the known filler phrases these models produce on silence; a match is treated as no speech rather than as input, and transcription runs at temperature zero to keep it literal. The documented limit is that this catches a filler phrase that is the whole transcript, not filler buried inside real speech. The decision was to fail toward no input rather than risk feeding the model a symptom the user never described.

Outcome

A technology demonstration. Not a medical device, and not medical advice. AI Doctor is live as a demo at aidoctor.techparser.io. A person describes a concern by voice, optionally attaches a photo, and gets a streamed, structured assessment with a Low, Moderate or High urgency badge and an optional spoken reply.

The design is what keeps the demo safe to run. Nothing is written to disk, so concurrent users on serverless hosting can never see each other's photos or recordings. Urgency is encoded in the first line of the answer, so a reader gets an instant signal while the rest streams in. Silent audio is treated as no speech rather than invented symptoms. And two providers per voice step can be compared on the same answer, so cost and quality are easy to weigh.

Questions about this project

What technology stack does AI Doctor use?
AI Doctor is a Next.js and React app in TypeScript. The chained voice and vision pipeline runs on Groq, using its vision model, Whisper for speech-to-text and Orpheus for speech, with ElevenLabs Scribe and TTS as the second provider behind a toggle for each voice step. Nothing is written to disk, so it runs safely on serverless hosting.
How does AI Doctor avoid turning silent audio into symptoms?
Speech models invent filler phrases on silence, which in a health tool would become fabricated symptoms. So transcripts are normalised and matched against the known filler phrases these models produce, and a match is treated as no speech rather than as input. Transcription also runs at temperature zero to stay literal. The documented limit is that it catches a filler phrase that is the whole transcript, not filler inside real speech.
Can Techparser build a multimodal AI assistant like AI Doctor for us?
Yes. Techparser built AI Doctor end to end: a stateless in-memory pipeline, chained speech, vision and voice models, a streamed urgency badge, swappable providers per step and a session gate that fails closed. The same approach fits any multimodal assistant that must stay responsive and leak nothing between users. Book a call to discuss your use case.