Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

ย 

History

24 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

EasySummarizer ๐Ÿ“š

EasySummarizer is a multimodal AI-powered document assistant that is developed to allow users to upload PDF documents or images and ask questions about their contents.

The application uses Retrieval-Augmented Generation (RAG) to retrieve relevant information from uploaded files and generate grounded responses using an LLM. It also supports OCR-based image processing and short-term conversational memory.

The application is deployed using Streamlit Community Cloud.


๐Ÿš€ Features

EasySummarizer provides the following features:

  • PDF document upload
  • Image upload
  • Image preview after upload
  • OCR-based text extraction from images
  • Automatic document chunking
  • HuggingFace-based text embeddings
  • Chroma vector database
  • Semantic similarity-based retrieval
  • LLM-powered question answering
  • Short-term conversation memory
  • Context-aware follow-up questions
  • Retrieved source/context display
  • Clear current document functionality
  • Separate PDF and image upload sections
  • Easy-to-use Streamlit interface
  • Streamlit Cloud deployment
  • Custom EasySummarizer branding

๐Ÿ—๏ธ Architecture

The application follows a modular RAG architecture.

                    EasySummarizer
                           โ”‚
              โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
              โ”‚                         โ”‚
         PDF Upload                Image Upload
              โ”‚                         โ”‚
        PyPDFLoader                 Tesseract
              โ”‚                         โ”‚
        Text Extraction             OCR Text
              โ”‚                         โ”‚
              โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                           โ”‚
                       Chunking
                           โ”‚
                    HuggingFace
                     Embeddings
                           โ”‚
                       Chroma DB
                           โ”‚
                    User Question
                           โ”‚
                 Short-Term Memory
                           โ”‚
                    Similarity Search
                           โ”‚
                    Retrieved Chunks
                           โ”‚
                Context + Chat History
                           โ”‚
                       Groq LLM
                           โ”‚
                         Answer

๐Ÿ”„ RAG Pipeline

The document processing pipeline follows these stages:

1. Document Upload

Users upload either a PDF document or an image through the Streamlit interface.

2. Text Extraction

For PDFs, PyPDFLoader is used to extract the document text.

For images, Tesseract OCR is used to extract textual information.

3. Chunking

The extracted text is divided into smaller overlapping chunks using RecursiveCharacterTextSplitter.

The following configuration is used:

Chunk size: 600
Chunk overlap: 100

4. Embeddings

Each text chunk is converted into a vector representation using:

sentence-transformers/all-MiniLM-L6-v2

5. Vector Storage

The generated embeddings are stored in a Chroma vector database.

Each uploaded file is assigned a unique Chroma collection to prevent different documents from being mixed together.

6. Retrieval

When a user asks a question, the question is embedded and compared against the stored vectors.

The system retrieves the top 4 most relevant chunks.

7. Generation

The retrieved chunks are provided as context to the Groq-hosted LLM.

The model is instructed to answer using the retrieved context and avoid inventing information that is not present in the uploaded content.


๐Ÿง  Short-Term Conversation Memory

EasySummarizer supports short-term conversational memory using Streamlit's session_state.

The application retains the most recent conversation messages and passes them to the LLM along with the retrieved document context.

This allows users to ask follow-up questions such as:

User:
What was the company's refund policy?

Assistant:
The company allowed refunds within 30 days...

User:
What were the exceptions?

Assistant:
The exceptions included...

The assistant is therefore able to use the previous conversation to understand references and follow-up questions.

The conversation history is cleared whenever the user uploads a new document or selects the clear-file option.


๐Ÿ–ผ๏ธ Image Processing

Images are processed using an OCR pipeline.

Image Upload
     โ†“
Temporary File
     โ†“
Tesseract OCR
     โ†“
Extracted Text
     โ†“
LangChain Document
     โ†“
Chunking
     โ†“
Embeddings
     โ†“
Chroma
     โ†“
Retrieval

The image itself is also displayed as a preview in the Streamlit application.

The OCR pipeline was particularly suitable for:

  • Screenshots
  • Receipts
  • Invoices
  • Forms
  • Scanned documents
  • Text-heavy images

๐Ÿ› ๏ธ Technology Stack

Technology Purpose
Python Application development
Streamlit Web application and UI
LangChain RAG pipeline orchestration
HuggingFace Text embeddings
Sentence Transformers Embedding model
Chroma Vector database
Groq LLM inference
Tesseract OCR Image text extraction
PyPDF PDF text extraction
GitHub Version control
Streamlit Community Cloud Deployment

๐Ÿ“ Project Structure

projectabc/
โ”‚
โ”œโ”€โ”€ assets/
โ”‚   โ””โ”€โ”€ logo.png
โ”‚
โ”œโ”€โ”€ app.py
โ”œโ”€โ”€ ingestion_pdf.py
โ”œโ”€โ”€ ingestion_images.py
โ”œโ”€โ”€ retrieval.py
โ”œโ”€โ”€ qa_chain.py
โ”‚
โ”œโ”€โ”€ requirements.txt
โ”œโ”€โ”€ packages.txt
โ”œโ”€โ”€ .gitignore
โ””โ”€โ”€ README.md

app.py

The main Streamlit application is responsible for:

  • UI rendering
  • File uploads
  • Image previews
  • Session state
  • Short-term chat memory
  • User interaction
  • Displaying retrieved context

ingestion_pdf.py

The PDF ingestion module is responsible for:

  • Processing uploaded PDFs
  • Extracting text
  • Chunking documents
  • Creating embeddings
  • Storing vectors in Chroma

ingestion_images.py

The image ingestion module is responsible for:

  • Processing uploaded images
  • Running Tesseract OCR
  • Extracting text
  • Chunking extracted content
  • Creating embeddings
  • Storing vectors in Chroma

retrieval.py

The retrieval module is responsible for:

  • Loading the appropriate Chroma collection
  • Loading the embedding model
  • Creating the retriever
  • Returning the most relevant document chunks

qa_chain.py

The question-answering module is responsible for:

  • Loading the Groq LLM
  • Constructing the RAG prompt
  • Combining retrieved context with conversation history
  • Generating the final response

๐Ÿ” Security

API credentials are kept outside the source code.

During local development, environment variables are used through .env.

For Streamlit Community Cloud deployment, the Groq API key is stored using Streamlit Secrets.

The .env file is excluded from GitHub using .gitignore.

Generated Chroma data and the local virtual environment are also excluded from version control.


โ˜๏ธ Deployment

The application is deployed using Streamlit Community Cloud.

The deployment requires:

requirements.txt

for Python dependencies and:

packages.txt

for system-level dependencies such as Tesseract OCR.

The packages.txt file contained:

tesseract-ocr

This allows the OCR functionality to run in the Linux-based Streamlit deployment environment.


โš™๏ธ Local Setup

The project is developed using a Python virtual environment.

The dependencies were installed using:

pip install -r requirements.txt

The application was started using:

streamlit run app.py

The application is then accessed through the local Streamlit URL.


๐Ÿ”‘ Environment Variables

For local development, the Groq API key is stored in a .env file:

GROQ_API_KEY="your_api_key"

The .env file is not committed to GitHub.

For Streamlit Cloud, the API key is added through the application's Secrets configuration.


๐ŸŽฏ Project Objective

EasySummarizer is developed to demonstrate an end-to-end implementation of a multimodal RAG application.

The project combines:

  • Document ingestion
  • OCR
  • Text chunking
  • Embeddings
  • Vector databases
  • Semantic retrieval
  • LLM generation
  • Conversational memory
  • Streamlit application development
  • Cloud deployment

The project demonstrates how unstructured user-provided content can be transformed into a searchable knowledge base and used to generate context-aware answers.


๐Ÿ”ฎ Future Improvements

The following improvements are identified for future versions:

  • Support for DOCX, PPTX and XLSX files
  • Vision-language model support for charts and diagrams
  • Improved semantic chunking
  • Hybrid keyword + vector retrieval
  • Reranking of retrieved chunks
  • Source citations with page numbers
  • Automatic document summarization
  • Downloadable summaries
  • Persistent user conversations
  • User authentication
  • Conversation history across sessions
  • Retrieval evaluation and monitoring
  • Improved OCR accuracy
  • Production-grade vector database deployment

๐Ÿ‘จโ€๐Ÿ’ป Developer

EasySummarizer is developed by AnandAnalytics - Ayush Anand.

๐ŸŒ anandanalytics.online


๐Ÿ“Œ Key Learning Outcomes

The project provides hands-on experience with the complete lifecycle of a modern RAG application:

Unstructured Data
       โ†“
Ingestion
       โ†“
Chunking
       โ†“
Embeddings
       โ†“
Vector Database
       โ†“
Retrieval
       โ†“
Context Construction
       โ†“
LLM
       โ†“
Conversational Response

It demonstrates how individual GenAI components can be combined into a complete, deployable AI application rather than using an LLM as a standalone chatbot.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages