Document AI

How Techparser Built AI PDF Chat, a Multi-Document RAG Assistant

Techparser built AI PDF Chat, a multi-document RAG assistant: answers cite the PDF and page, with per-user vector isolation and background ingestion.

Client
Techparser product build
Industry
Document AI
Timeline
Built, available for deployment
Team
1 senior full-stack engineer

Want results like these?




Results

SSE
Answers stream as they are written
Per user
Every vector search filtered by the signed-in user
Cited
Each answer cites the PDF and page it used

The problem

A signed-in user uploads several PDFs and asks questions across all of them or inside one. Every answer cites the PDF and page, and the citation opens the document with the passage highlighted. The embedding provider's free tier rate-limits hard, and through the library in use it returns empty results instead of an error when it does. Large PDFs produce many chunks, so processing cannot happen inside the upload request. Many users share one vector store, and the model provider had to be swappable.

The failure modes are quiet ones. Embedding during the upload would time out on big files. A missed rate limit would silently fill the index with empty vectors and degrade every answer. A missing user filter would show one person's documents to another. And vectors left behind after a delete would keep surfacing content the user thought was gone.

What Techparser built

  • The upload responds immediately and each file is extracted, chunked and embedded by a background queue, one file at a time, with its status shown in the app; one file failing never holds up the rest.
  • On startup, files that were mid-processing are put back in the queue, so a restart never leaves a document stuck.
  • Chat and embeddings are chosen from configuration, with a hosted provider by default and a local model as an option, so the vendor can change without touching the pipeline.
  • Empty or short vectors are treated as a hidden rate-limit response and retried with exponential backoff; batches are paced to stay under the limit, and if it never recovers the user sees a plain message.
  • Each re-ingest clears the file's old vectors first, so retries never create duplicates, and a scanned PDF with no text is marked failed with a useful message rather than shown as ready.
  • Searches always filter by the signed-in user, and by the file in single-document mode, on indexed fields, with a relevance threshold so weak matches are never passed to the model.
  • Viewing, deleting and retrying files, and reading chat history, are all scoped to the signed-in user, and sign-in endpoints are rate-limited separately.
  • Deleting a PDF removes the file, its record and its vectors, and any conversation that cited only that file, while keeping conversations that drew on other documents too.
  • Answers stream over SSE as they are written and are restricted to the retrieved content; citations are grouped per file and page and carry the best-matching passage, which the viewer highlights.

Decisions that mattered

Ingestion off the request path

Large PDFs produce many chunks, so embedding them inside the upload request would time out. The upload instead responds immediately and hands each file to a background queue that extracts, chunks and embeds one file at a time, with a visible status of queued, processing, ready or failed. One file failing never blocks the rest, and on startup any file left mid-processing is put back in the queue so a restart never strands a document. The trade-off is that a document is not searchable the instant it is uploaded, in exchange for an upload that never times out and work that always resumes.

Designing around a silent rate limit

The embedding provider's free tier rate-limits hard, and through the library in use it returns empty results instead of an error. Left alone, that would quietly fill the index with empty vectors and degrade every answer. So empty or short vectors are treated as a hidden rate-limit response and retried with exponential backoff, and batches are paced to stay under the limit in the first place. If it never recovers, the user sees a plain message about the limit rather than silent, broken results. The decision was to assume the provider lies about failure and verify every vector before trusting it.

Isolation and deletes that leave nothing

Many users share one vector store, so isolation cannot be assumed. Every search filters by the signed-in user, and by the file in single-document mode, on indexed fields, with a relevance threshold so weak matches never reach the model. Viewing, deleting and retrying files are all scoped to the owner. Deleting a PDF removes the file, its record, its vectors and any conversation that cited only that file, while keeping conversations that also drew on other documents. The decision was to make a delete truly complete, so content a user removed can never resurface in a later answer.

Outcome

AI PDF Chat is built and available for deployment. A signed-in user uploads several PDFs and asks questions across all of them or inside one, and every answer cites the PDF and page, with the citation opening the document at the highlighted passage.

The design is what makes the assistant dependable at scale. Ingestion runs off the request path, so large uploads never time out and a restart resumes where it left off. Every search is filtered by user, so one person's documents never appear in another's answers. Because the provider is chosen from configuration, the chat and embedding vendors can change without touching the pipeline, and because deletes remove every vector, removed content never resurfaces.

Questions about this project

What technology stack does AI PDF Chat use?
The app is a Next.js frontend with an Express backend. Retrieval is built on LangChain with a Qdrant vector store, Gemini for embeddings and Groq's Llama for chat, both chosen from configuration so the vendors can be swapped. Zod validates every input, and the whole stack runs locally under Docker Compose.
How does AI PDF Chat handle an embedding provider that silently rate-limits?
The provider's free tier returns empty results instead of an error when it is throttled, which would quietly poison the index. So empty or short vectors are treated as a hidden rate-limit response and retried with exponential backoff, and batches are paced to stay under the limit. If it never recovers, the user gets a plain message instead of broken answers, and each re-ingest clears old vectors first so retries never duplicate.
Can Techparser build a multi-document RAG assistant like AI PDF Chat for us?
Yes. Techparser built AI PDF Chat end to end: background ingestion, per-user vector isolation, citations that open the source page, streamed answers and a swappable model provider. The same pattern fits any document-heavy workflow where answers must be grounded in your own files and cite where they came from. Book a call to talk through your documents.