← Selected engineering work

Personal project · Working prototype · AI / RAG

AskBase: answers you can check

I built AskBase to ask questions about PDFs and websites without losing the link to the original information. It brings together content import, search, streamed answers, and source citations in one app.

Next.js · TypeScript · Supabase PostgreSQL · pgvector · OpenAI · Vercel AI SDK

The problem

A confident AI answer is hard to trust if you cannot see where it came from. My goal was to make the answer and its evidence easy to read together, while keeping the underlying retrieval flow understandable.

How it works

  1. Bring in content

    Extract text from a PDF or import a website. The crawler follows links on the same site, with a default limit of 25 pages, and reports progress as it works.

  2. Make it searchable

    Split text into chunks, embed them with OpenAI, and store the vectors alongside their source details in Supabase PostgreSQL. PDF embeddings are sent in batches of 64.

  3. Find useful passages

    Embed the question and retrieve four chunks using pgvector. Ranking combines vector similarity with a small adjustment based on earlier feedback.

  4. Answer and show sources

    Pass the full retrieved passages to GPT-4o-mini through the Vercel AI SDK. Stream the answer, ask the model to cite numbered sources, and let readers inspect source snippets or open the original URL.

Existing project recording. The live interface may differ.

Engineering choices

Keep retrieval easy to inspect

I used PostgreSQL and pgvector so document records, vectors, queries, and feedback could live together. The retrieval logic is a SQL function rather than a separate search service.

Make the wait visible

Website import reports progress, and chat streams as the answer is generated. These make longer operations easier to follow without hiding the work behind a spinner.

Go beyond a chat box

I added answer feedback and an admin view for query history, feedback rates, frequently used chunks, and possible knowledge gaps. These provide clues about where an answer went wrong.

Preserve work during a migration

When chat moved to the Vercel AI SDK, I added a normalizer for older locally stored messages. It keeps existing conversations readable alongside the new message format.

What works, and what I’d improve

The prototype connects the complete document Q&A flow and includes tools to inspect answers and feedback. I have not measured an accuracy or latency improvement against a benchmark.

  • Measure retrieval quality. Add questions with known source passages, then compare vector search with keyword search and reranking. Feedback scores alone do not prove an answer is correct.
  • Verify citations. Sources are selected by retrieval and referenced by the model. A citation still needs checking against the actual claim.
  • Handle follow-ups better. Chat passes conversation history to the model, but retrieval embeds only the latest user message. Rewrite follow-up questions before searching.
  • Prepare for private use. Add authentication, enforce document access in retrieval, protect admin routes, and fail clearly when dependencies are missing. The current demo is not a private document workspace.

What this project shows

I can take an AI product beyond its prompt: build the interface, connect ingestion and retrieval, handle streaming and persistence, and make failures easier to investigate. I also know where a prototype needs stronger evaluation and access controls.

See my experience and get in touch →