DOC-Analyser-Using-LLM

๐Ÿ“„ AI Document Analyzer

Python Streamlit Status License

AI-powered Document Analysis Web Application
Leverage the power of Large Language Models (LLMs) to instantly extract insights and answer questions from PDF and TXT documents.


๐Ÿš€ Overview

The AI Document Analyzer is a Streamlit-based web app that allows users to upload documents and interact with their content using cutting-edge AI models.
Whether you want to analyze reports, research papers, or any text-based content, this tool provides real-time, context-aware answers.


โœจ Features


๐Ÿ“„ Supported File Types

File Type Extraction Method
PDF PyPDF2 library for structured text extraction
TXT Direct UTF-8 file reading

โš ๏ธ Note: Scanned or image-based PDFs may not extract text correctly.


๐Ÿค– AI Models

1๏ธโƒฃ Ollama (Offline)

2๏ธโƒฃ Gemini (Free Cloud)


๐Ÿ›  Installation

Prerequisites

Install Dependencies

pip install streamlit PyPDF2 google-genai

Optional: Ollama Setup

Install Ollama: https://ollama.ai/

Pull the required model:

ollama pull tinyllama

Start Ollama service:

ollama serve

Optional: Gemini Setup

Get a free API key: https://makersuite.google.com/app/apikey

Set environment variable:

# Windows
set GEMINI_API_KEY=your_api_key_here

# Linux/Mac
export GEMINI_API_KEY=your_api_key_here

๐Ÿ’ก Usage

Start the app:

streamlit run app.py

Open browser at http://localhost:8501

  1. Upload PDF or TXT document
  2. Select AI backend:
    • Ollama (Offline) โ†’ Local processing
    • Gemini (Free Cloud) โ†’ Cloud-based processing
  3. Enter your question about the document
  4. Click Analyze Document to get AI-powered answers

๐Ÿ— Project Structure

doc-analysier/
โ”œโ”€โ”€ app.py              # Main Streamlit app
โ”œโ”€โ”€ README.md           # Documentation
โ””โ”€โ”€ .git/               # Git repository files

๐Ÿ”‘ Core Functions

Function Description
extract_text_from_file() Extract text from PDF/TXT
ask_ollama() Query local Ollama LLM
ask_gemini() Query Google Gemini API
main() Handles Streamlit UI and workflow

โš ๏ธ Limitations


๐Ÿ›  Troubleshooting

Issue Solution
โ€œPyPDF2 not installedโ€ Run pip install PyPDF2
โ€œOllama not foundโ€ Install Ollama and start the service
โ€œGEMINI_API_KEY not foundโ€ Set environment variable correctly
โ€œUnsupported file typeโ€ Upload PDF or TXT only
โ€œPDF text extraction failsโ€ Check if PDF is image-based or complex

๐Ÿ— Development Stack


๐Ÿ“ Future Enhancements


๐Ÿ“œ License

MIT License โ€“ Open source and free to use, modify, and redistribute.