A modern Retrieval-Augmented Generation (RAG) application that allows users to upload one or more PDF documents and ask questions using a local Large Language Model (LLM) powered by Ollama and Qdrant Vector Database.
โญ If you find this project useful, consider giving it a star!
- Overview
- Features
- System Architecture
- Project Structure
- Technology Stack
- Installation
- Configuration
- Usage
- Screenshots
- RAG Workflow
- Future Roadmap
- Contributing
- License
- Author
AI PDF Assistant is a local Retrieval-Augmented Generation (RAG) application that enables users to interact with PDF documents using natural language.
Instead of manually searching through long PDF files, users can upload one or more documents and ask questions in plain English. The application retrieves the most relevant information using semantic search and generates answers with a local Large Language Model.
Unlike cloud-based AI applications, this project runs entirely on your machine using Ollama and Qdrant, helping keep your documents private while providing fast responses.
The project is designed as a modular, scalable codebase suitable for learning, portfolio projects, and further development.
- ๐ค Upload one or multiple PDF documents
- ๐ Automatic PDF text extraction
- โ๏ธ Intelligent text chunking
- ๐ง Metadata-aware document processing
- ๐๏ธ Document Manager
- ๐๏ธ Delete individual documents
- ๐งน Clear the complete knowledge base
- ๐ฌ Interactive chat interface
- ๐ง Conversation memory
- ๐ Context-aware responses
- ๐ Rich source references
- ๐ Semantic document search
- โก Fast local inference using Ollama
- ๐ Chat history
- Automatic text chunking
- Sentence embeddings
- Vector similarity search
- Context retrieval
- Metadata filtering
- Grounded AI responses
- Multi-document search
- ๐จ Modern Streamlit interface
- ๐ Sidebar document manager
- ๐ Knowledge base statistics
- ๐ค Multiple PDF upload
- ๐งน Clear Chat
- ๐๏ธ Clear Database
- ๐ฑ Simple and responsive layout
โ 100% Local AI
โ No OpenAI API Required
โ Privacy Friendly
โ Multiple PDF Support
โ Conversation Memory
โ Rich Source References
โ Qdrant Vector Database
โ Semantic Search
โ Modular Architecture
โ Production-Oriented Code Structure
| Category | Technology |
|---|---|
| Programming Language | Python 3.11+ |
| User Interface | Streamlit |
| Local LLM | Ollama (Llama 3.2) |
| Vector Database | Qdrant |
| Embedding Model | all-MiniLM-L6-v2 |
| PDF Processing | PyMuPDF |
| Vector Search | Semantic Similarity Search |
| Version Control | Git & GitHub |
| Dependency Management | uv |
This project demonstrates the implementation of a complete Retrieval-Augmented Generation (RAG) pipeline using modern AI tools and frameworks.
It combines document processing, semantic search, vector databases, and local large language models into a single application capable of answering questions based on uploaded PDF documents.
The project is designed with a modular architecture, making it easy to extend with additional features such as streaming responses, PDF previews, dashboards, hybrid search, and cloud deployment.
This application can be used for:
- ๐ Study Notes
- ๐ Research Papers
- ๐ Company Documentation
- ๐ User Manuals
- ๐ E-books
- ๐งพ Technical Documentation
- ๐ Educational Material
- ๐ Project Reports
Version: v1.0
- โ Multiple PDF Upload
- โ Local LLM (Ollama)
- โ Semantic Search
- โ Qdrant Vector Database
- โ Conversation Memory
- โ Rich Source References
- โ Document Manager
- โ Sidebar Statistics
- โ Clear Database
- โ Modular Service Architecture
The primary goals of this project are:
- Build a complete local RAG application
- Learn vector databases and semantic search
- Explore local LLM deployment with Ollama
- Develop a modular and scalable architecture
- Create a portfolio-ready AI application
- Demonstrate modern AI engineering practices
The AI PDF Assistant follows a modular Retrieval-Augmented Generation (RAG) architecture.
+----------------------+
| Upload PDF(s) |
+----------+-----------+
|
โผ
+----------------------+
| PDF Text Loader |
+----------+-----------+
|
โผ
+----------------------+
| Text Chunking |
+----------+-----------+
|
โผ
+----------------------+
| Sentence Embeddings |
+----------+-----------+
|
โผ
+----------------------+
| Qdrant Vector DB |
+----------+-----------+
โฒ
|
User Question
|
โผ
+----------------------+
| Semantic Retrieval |
+----------+-----------+
|
โผ
+----------------------+
| Ollama LLM |
+----------+-----------+
|
โผ
AI Answer + Sources
The application follows the Retrieval-Augmented Generation workflow.
The user uploads one or more PDF documents using the Streamlit interface.
โ
The application extracts readable text from every page of the uploaded PDF.
โ
Large documents are divided into smaller chunks so they can be embedded efficiently.
โ
Each text chunk is converted into a high-dimensional vector using the embedding model.
โ
The generated vectors are stored inside Qdrant together with useful metadata.
Example metadata:
{
"document_id": "...",
"filename": "AI.pdf",
"chunk_index": 12,
"text": "Artificial Intelligence..."
}โ
The user asks a question in natural language.
โ
The question is embedded and compared against all stored vectors.
The most relevant chunks are retrieved.
โ
Retrieved chunks are combined into a prompt.
Conversation history is also included.
โ
The prompt is sent to the local Llama model.
โ
The assistant generates an answer together with supporting source references.
AI PDF Assistant
โ
โโโ app
โ โโโ api
โ โ โโโ upload.py
โ โ
โ โโโ core
โ โ โโโ config.py
โ โ โโโ constants.py
โ โ โโโ logger.py
โ โ
โ โโโ services
โ โ โโโ chunker.py
โ โ โโโ database_service.py
โ โ โโโ document_service.py
โ โ โโโ embedding_service.py
โ โ โโโ ingestion_service.py
โ โ โโโ llm_service.py
โ โ โโโ pdf_loader.py
โ โ โโโ search_service.py
โ โ โโโ vector_service.py
โ โ
โ โโโ main.py
โ
โโโ frontend
โ โโโ chat_page.py
โ โโโ sidebar.py
โ โโโ streamlit_app.py
โ โโโ upload_page.py
โ
โโโ tests
โ
โโโ pyproject.toml
โโโ uv.lock
โโโ README.md
โโโ LICENSE
The application is divided into reusable service modules.
| Service | Responsibility |
|---|---|
| PDF Loader | Extract text from PDF files |
| Chunker | Split documents into chunks |
| Embedding Service | Generate vector embeddings |
| Vector Service | Store and search vectors in Qdrant |
| Search Service | Retrieve relevant chunks |
| LLM Service | Communicate with Ollama |
| Database Service | Clear and manage the knowledge base |
| Document Service | Manage uploaded documents |
Each document chunk is stored with metadata to support advanced features.
Current metadata includes:
- Document ID
- Filename
- Source Path
- Chunk Index
- Chunk Text
This metadata enables:
- Individual document deletion
- Rich source references
- Document statistics
- Future page-aware citations
PDF
โ
โผ
Text Extraction
โ
โผ
Chunking
โ
โผ
Embeddings
โ
โผ
Qdrant
โ
โผ
Semantic Search
โ
โผ
Context
โ
โผ
Ollama
โ
โผ
Answer + Sources
The project follows several software engineering principles:
- Modular architecture
- Separation of concerns
- Service-oriented design
- Reusable components
- Scalable project structure
- Easy future extensibility
- Clean and maintainable code
Follow these steps to set up the project on your local machine.
Before running the project, make sure you have the following installed:
- Python 3.11 or later
- Git
- Ollama
- Docker (for Qdrant)
- uv (Python package manager)
git clone https://github.com/sagarsoni7254-eng/ai-pdf-assistant.gitMove into the project directory:
cd ai-pdf-assistantpython -m venv .venvActivate:
.venv\Scripts\activatepython3 -m venv .venvActivate:
source .venv/bin/activateThis project uses uv for dependency management.
Install all required packages:
uv syncThis command installs all dependencies defined in:
- pyproject.toml
- uv.lock
Download Ollama from:
Verify installation:
ollama --versionDownload the required model:
ollama pull llama3.2Start the Ollama server:
ollama serveRun Qdrant locally using Docker:
docker run -p 6333:6333 qdrant/qdrantVerify Qdrant is running:
Open your browser:
http://localhost:6333/dashboard
Launch Streamlit:
streamlit run frontend/streamlit_app.pyOpen:
http://localhost:8501
The application should now be running successfully.
Current default configuration:
| Setting | Value |
|---|---|
| LLM | llama3.2 |
| Embedding Model | all-MiniLM-L6-v2 |
| Vector Database | Qdrant |
| Frontend | Streamlit |
| Language | Python |
Upload one or more PDF documents.
The application automatically:
- Extracts text
- Splits documents into chunks
- Generates embeddings
- Stores vectors in Qdrant
Ask questions using natural language.
Examples:
What is Artificial Intelligence?
Explain Machine Learning.
Summarize Chapter 5.
Compare CNN and RNN.
List important interview questions.
Explain this topic for beginners.
The assistant:
- Retrieves relevant document chunks
- Builds context
- Sends the prompt to Ollama
- Generates an answer
- Displays source references
Start Ollama:
ollama serveDownload the model:
ollama pull llama3.2Make sure Docker is running.
Restart Qdrant:
docker run -p 6333:6333 qdrant/qdrantActivate the virtual environment:
source .venv/bin/activateThen install dependencies again:
uv sync- โ Multiple PDF Upload
- โ Semantic Search
- โ Conversation Memory
- โ Rich Source References
- โ Local LLM
- โ Qdrant Vector Database
- โ Modular Architecture
- โ Modern UI
- โ Easy Installation
Screenshots will be updated as the project evolves.
The main interface where users can upload documents and interact with the AI assistant.
screenshots/home.png
Upload one or multiple PDF documents.
screenshots/upload.png
Ask questions about uploaded documents using natural language.
screenshots/chat.png
Every response includes supporting document references.
screenshots/sources.png
Manage uploaded documents and maintain your knowledge base.
screenshots/sidebar.png
The following features are planned for future releases.
- โก Streaming AI Responses
- ๐ Page-aware Source References
- ๐ Knowledge Base Dashboard
- ๐ Similarity Score Display
- ๐ PDF Preview
- ๐ค Export Chat (PDF & Markdown)
- ๐จ Improved User Interface
- ๐ Dark / Light Theme
- ๐ Multi-user Authentication
- โ Cloud Deployment
- ๐ณ Docker Compose Support
- ๐ Hybrid Search (Semantic + Keyword)
- ๐ DOCX & TXT Support
- ๐ Voice Input
- ๐ Text-to-Speech Responses
Contributions are welcome!
If you'd like to improve this project:
- Fork the repository.
- Create a new branch.
git checkout -b feature/my-feature- Commit your changes.
git commit -m "Add my feature"- Push your branch.
git push origin feature/my-feature- Open a Pull Request.
If you find a bug or want to request a feature, please create a GitHub Issue.
Please include:
- Operating System
- Python Version
- Error Message
- Steps to Reproduce
This project demonstrates practical experience with:
- Python Development
- Streamlit Applications
- Retrieval-Augmented Generation (RAG)
- Semantic Search
- Vector Databases
- Local Large Language Models
- Software Architecture
- Git & GitHub
- AI Application Development
This project is licensed under the MIT License.
See the LICENSE file for details.
AI & Machine Learning Enthusiast
Interested in:
- ๐ค Artificial Intelligence
- ๐ง Machine Learning
- ๐ Retrieval-Augmented Generation (RAG)
- ๐ฌ Large Language Models
- ๐ Python Development
GitHub:
https://github.com/sagarsoni7254-eng
This project would not have been possible without the amazing open-source community.
Special thanks to:
- Python
- Streamlit
- Ollama
- Qdrant
- Sentence Transformers
- PyMuPDF
- GitHub
If you found this project useful:
โญ Star this repository
๐ด Fork the project
๐ฌ Share your feedback
Your support motivates future improvements.