Early-stage open-source project

Ask your documents.
See the evidence.

DocAI is a document-intelligence assistant for PDF and TXT files. It retrieves relevant passages, generates concise answers from those sources, and stores citations so you can inspect the evidence behind every response.

PDF + TXT Grounded answers Stored citations
DocAI workspace
Source-grounded
research-notes.pdf 18 pages · processed
You What are the main limitations described in this document?
AI Grounded answer

The document identifies three main limitations: restricted input coverage, dependence on text quality, and the need for human review in high-stakes settings. [1] [2]

1 research-notes.pdf · passage 12
2 research-notes.pdf · passage 17
7 MB upload limit
PDF / TXT current file support
User-scoped document access
Inspectable stored citations
Built for evidence-aware Q&A

Useful answers without hiding the source.

DocAI keeps the retrieval path understandable: uploaded documents are processed into searchable passages, relevant passages are selected, and the answer is generated from those sources.

Grounded question answering

The model receives selected source passages instead of the entire document and is instructed to answer from those passages.

Persistent citations

Answers keep references to the source document, passage, filename, and excerpt so the evidence trail remains available for review.

Transparent retrieval

Current retrieval is deliberately lexical and explainable, using direct terms, related terms, title matches, and exact-phrase bonuses.

User-scoped data

Protected procedures check the signed-in user before accessing documents, conversations, messages, and citations.

Controlled ingestion

PDF and TXT uploads are validated by extension, MIME type, file size, and readable text before processing.

Full-stack foundation

React, TypeScript, tRPC, Express, MySQL/Drizzle, object storage, and automated tests form the current implementation.

How it works

From upload to cited answer.

DocAI is designed to keep the path between a document and an answer explicit.

  1. 1

    Upload a document

    An authenticated user uploads a text-based PDF or TXT file. Unsupported or unreadable files are rejected.

  2. 2

    Extract and chunk text

    Readable text is normalized and divided into overlapping chunks for search and retrieval.

  3. 3

    Retrieve relevant passages

    The current ranking system scores lexical signals to select passages related to the user's question.

  4. 4

    Generate from the sources

    Only the selected passages are sent to the configured language-model integration with grounding instructions.

  5. 5

    Save the evidence trail

    The response, conversation history, and citations are persisted so users can return to the answer and inspect its sources.

Project principles

Grounded by design. Clear about limitations.

Evidence over opacity

A generated answer should make it easier—not harder—to inspect the material it relies on.

Explicit boundaries

A citation does not make an answer automatically correct. Users should review original documents and cited excerpts.

Practical iteration

DocAI is in active development. Current limitations include no OCR for scanned image-only PDFs and lexical rather than embedding-based retrieval.

Open source · In development

Explore the implementation.

Read the code, architecture notes, tests, and current limitations in the public repository.