EasySummarizer is a multimodal AI-powered document assistant that is developed to allow users to upload PDF documents or images and ask questions about their contents.
The application uses Retrieval-Augmented Generation (RAG) to retrieve relevant information from uploaded files and generate grounded responses using an LLM. It also supports OCR-based image processing and short-term conversational memory.
The application is deployed using Streamlit Community Cloud.
EasySummarizer provides the following features:
- PDF document upload
- Image upload
- Image preview after upload
- OCR-based text extraction from images
- Automatic document chunking
- HuggingFace-based text embeddings
- Chroma vector database
- Semantic similarity-based retrieval
- LLM-powered question answering
- Short-term conversation memory
- Context-aware follow-up questions
- Retrieved source/context display
- Clear current document functionality
- Separate PDF and image upload sections
- Easy-to-use Streamlit interface
- Streamlit Cloud deployment
- Custom EasySummarizer branding
The application follows a modular RAG architecture.
EasySummarizer
โ
โโโโโโโโโโโโโโดโโโโโโโโโโโโโ
โ โ
PDF Upload Image Upload
โ โ
PyPDFLoader Tesseract
โ โ
Text Extraction OCR Text
โ โ
โโโโโโโโโโโโโโฌโโโโโโโโโโโโโ
โ
Chunking
โ
HuggingFace
Embeddings
โ
Chroma DB
โ
User Question
โ
Short-Term Memory
โ
Similarity Search
โ
Retrieved Chunks
โ
Context + Chat History
โ
Groq LLM
โ
Answer
The document processing pipeline follows these stages:
Users upload either a PDF document or an image through the Streamlit interface.
For PDFs, PyPDFLoader is used to extract the document text.
For images, Tesseract OCR is used to extract textual information.
The extracted text is divided into smaller overlapping chunks using RecursiveCharacterTextSplitter.
The following configuration is used:
Chunk size: 600
Chunk overlap: 100
Each text chunk is converted into a vector representation using:
sentence-transformers/all-MiniLM-L6-v2
The generated embeddings are stored in a Chroma vector database.
Each uploaded file is assigned a unique Chroma collection to prevent different documents from being mixed together.
When a user asks a question, the question is embedded and compared against the stored vectors.
The system retrieves the top 4 most relevant chunks.
The retrieved chunks are provided as context to the Groq-hosted LLM.
The model is instructed to answer using the retrieved context and avoid inventing information that is not present in the uploaded content.
EasySummarizer supports short-term conversational memory using Streamlit's session_state.
The application retains the most recent conversation messages and passes them to the LLM along with the retrieved document context.
This allows users to ask follow-up questions such as:
User:
What was the company's refund policy?
Assistant:
The company allowed refunds within 30 days...
User:
What were the exceptions?
Assistant:
The exceptions included...
The assistant is therefore able to use the previous conversation to understand references and follow-up questions.
The conversation history is cleared whenever the user uploads a new document or selects the clear-file option.
Images are processed using an OCR pipeline.
Image Upload
โ
Temporary File
โ
Tesseract OCR
โ
Extracted Text
โ
LangChain Document
โ
Chunking
โ
Embeddings
โ
Chroma
โ
Retrieval
The image itself is also displayed as a preview in the Streamlit application.
The OCR pipeline was particularly suitable for:
- Screenshots
- Receipts
- Invoices
- Forms
- Scanned documents
- Text-heavy images
| Technology | Purpose |
|---|---|
| Python | Application development |
| Streamlit | Web application and UI |
| LangChain | RAG pipeline orchestration |
| HuggingFace | Text embeddings |
| Sentence Transformers | Embedding model |
| Chroma | Vector database |
| Groq | LLM inference |
| Tesseract OCR | Image text extraction |
| PyPDF | PDF text extraction |
| GitHub | Version control |
| Streamlit Community Cloud | Deployment |
projectabc/
โ
โโโ assets/
โ โโโ logo.png
โ
โโโ app.py
โโโ ingestion_pdf.py
โโโ ingestion_images.py
โโโ retrieval.py
โโโ qa_chain.py
โ
โโโ requirements.txt
โโโ packages.txt
โโโ .gitignore
โโโ README.md
The main Streamlit application is responsible for:
- UI rendering
- File uploads
- Image previews
- Session state
- Short-term chat memory
- User interaction
- Displaying retrieved context
The PDF ingestion module is responsible for:
- Processing uploaded PDFs
- Extracting text
- Chunking documents
- Creating embeddings
- Storing vectors in Chroma
The image ingestion module is responsible for:
- Processing uploaded images
- Running Tesseract OCR
- Extracting text
- Chunking extracted content
- Creating embeddings
- Storing vectors in Chroma
The retrieval module is responsible for:
- Loading the appropriate Chroma collection
- Loading the embedding model
- Creating the retriever
- Returning the most relevant document chunks
The question-answering module is responsible for:
- Loading the Groq LLM
- Constructing the RAG prompt
- Combining retrieved context with conversation history
- Generating the final response
API credentials are kept outside the source code.
During local development, environment variables are used through .env.
For Streamlit Community Cloud deployment, the Groq API key is stored using Streamlit Secrets.
The .env file is excluded from GitHub using .gitignore.
Generated Chroma data and the local virtual environment are also excluded from version control.
The application is deployed using Streamlit Community Cloud.
The deployment requires:
requirements.txt
for Python dependencies and:
packages.txt
for system-level dependencies such as Tesseract OCR.
The packages.txt file contained:
tesseract-ocr
This allows the OCR functionality to run in the Linux-based Streamlit deployment environment.
The project is developed using a Python virtual environment.
The dependencies were installed using:
pip install -r requirements.txtThe application was started using:
streamlit run app.pyThe application is then accessed through the local Streamlit URL.
For local development, the Groq API key is stored in a .env file:
GROQ_API_KEY="your_api_key"The .env file is not committed to GitHub.
For Streamlit Cloud, the API key is added through the application's Secrets configuration.
EasySummarizer is developed to demonstrate an end-to-end implementation of a multimodal RAG application.
The project combines:
- Document ingestion
- OCR
- Text chunking
- Embeddings
- Vector databases
- Semantic retrieval
- LLM generation
- Conversational memory
- Streamlit application development
- Cloud deployment
The project demonstrates how unstructured user-provided content can be transformed into a searchable knowledge base and used to generate context-aware answers.
The following improvements are identified for future versions:
- Support for DOCX, PPTX and XLSX files
- Vision-language model support for charts and diagrams
- Improved semantic chunking
- Hybrid keyword + vector retrieval
- Reranking of retrieved chunks
- Source citations with page numbers
- Automatic document summarization
- Downloadable summaries
- Persistent user conversations
- User authentication
- Conversation history across sessions
- Retrieval evaluation and monitoring
- Improved OCR accuracy
- Production-grade vector database deployment
EasySummarizer is developed by AnandAnalytics - Ayush Anand.
The project provides hands-on experience with the complete lifecycle of a modern RAG application:
Unstructured Data
โ
Ingestion
โ
Chunking
โ
Embeddings
โ
Vector Database
โ
Retrieval
โ
Context Construction
โ
LLM
โ
Conversational Response
It demonstrates how individual GenAI components can be combined into a complete, deployable AI application rather than using an LLM as a standalone chatbot.