A RAG-based document analysis application that allows users to upload a PDF, ask questions about its content, and receive grounded answers based on relevant document sections with source references and page numbers.
- 📄 Upload and extract text from PDF documents while preserving page numbers
- ✂️ Split documents into smaller overlapping chunks
- 🔢 Generate semantic embeddings using Sentence Transformers
- 🔎 Perform semantic similarity search using FAISS
- 🤖 Generate grounded answers using the Google Gemini API
- 📌 Display relevant document excerpts with source page references
- 🛡️ Reduce unsupported answers by grounding responses in retrieved document context
- 🔍 Analyze documents to identify important sections and potentially notable clauses
- Python
- Streamlit
- Google Gemini API
- Sentence Transformers
- FAISS
- PyMuPDF
rag-document-explainer/
│
├── app.py
├── .env
├── .gitignore
├── README.md
├── requirements.txt
└── utils/
├── pdf_reader.py
├── chunker.py
├── vector_store.py
└── gemini_client.py
- Support for multiple document uploads
- Conversational memory across questions
- Hybrid keyword + semantic retrieval
- Retrieval quality evaluation
- Support for additional document formats
- Improved document-level analysis
- Advanced retrieval techniques such as reranking