AI-powered Document Analysis Web Application
Leverage the power of Large Language Models (LLMs) to instantly extract insights and answer questions from PDF and TXT documents.
The AI Document Analyzer is a Streamlit-based web app that allows users to upload documents and interact with their content using cutting-edge AI models.
Whether you want to analyze reports, research papers, or any text-based content, this tool provides real-time, context-aware answers.
| File Type | Extraction Method |
|---|---|
| PyPDF2 library for structured text extraction | |
| TXT | Direct UTF-8 file reading |
โ ๏ธ Note: Scanned or image-based PDFs may not extract text correctly.
tinyllama modelgemini-flash-latest)pip package managerpip install streamlit PyPDF2 google-genai
Install Ollama: https://ollama.ai/
Pull the required model:
ollama pull tinyllama
Start Ollama service:
ollama serve
Get a free API key: https://makersuite.google.com/app/apikey
Set environment variable:
# Windows
set GEMINI_API_KEY=your_api_key_here
# Linux/Mac
export GEMINI_API_KEY=your_api_key_here
Start the app:
streamlit run app.py
Open browser at http://localhost:8501
doc-analysier/
โโโ app.py # Main Streamlit app
โโโ README.md # Documentation
โโโ .git/ # Git repository files
| Function | Description |
|---|---|
extract_text_from_file() |
Extract text from PDF/TXT |
ask_ollama() |
Query local Ollama LLM |
ask_gemini() |
Query Google Gemini API |
main() |
Handles Streamlit UI and workflow |
| Issue | Solution |
|---|---|
| โPyPDF2 not installedโ | Run pip install PyPDF2 |
| โOllama not foundโ | Install Ollama and start the service |
| โGEMINI_API_KEY not foundโ | Set environment variable correctly |
| โUnsupported file typeโ | Upload PDF or TXT only |
| โPDF text extraction failsโ | Check if PDF is image-based or complex |
MIT License โ Open source and free to use, modify, and redistribute.