Skip to content

Repository files navigation

Handwritten Text Transcription App

A fully local desktop application that processes screenshots and PDF documents to recognize and transcribe handwritten text, marked options, and learn from user-provided examples to imitate writing styles.

Features

  • Multi-format Input: Accept screenshots (PNG, JPG) and PDF documents
  • Handwritten Text Recognition: Advanced OCR capabilities optimized for handwriting
  • Marked Options Recognition: Detect and transcribe checkboxes, radio buttons, and form selections
  • Writing Style Learning: Machine learning to adapt to user's handwriting patterns
  • Selective PDF Processing: Specify individual pages or page ranges for processing
  • Full Screenshot Processing: Process entire screenshot images
  • Local Processing: No cloud dependencies, ensuring complete data privacy

Technology Stack

  • Desktop Framework: Electron with React and TypeScript
  • Backend Processing: Python with FastAPI for ML/OCR processing
  • OCR Engine: Tesseract OCR with custom handwriting models
  • ML Framework: TensorFlow for style learning and adaptation
  • Image Processing: OpenCV for preprocessing and enhancement
  • PDF Processing: PyMuPDF for PDF manipulation
  • Database: SQLite for local data storage

Installation

Prerequisites

  1. Node.js (v16 or later)
  2. Python (v3.8 or later)
  3. Tesseract OCR installed on your system

Installing Tesseract OCR

Windows:

# Download and install from: https://github.com/UB-Mannheim/tesseract/wiki
# Or using chocolatey:
choco install tesseract

macOS:

brew install tesseract

Linux (Ubuntu/Debian):

sudo apt update
sudo apt install tesseract-ocr

Setup

  1. Clone the repository:
git clone <repository-url>
cd handwriting-transcription-app
  1. Install Node.js dependencies:
npm install
  1. Set up Python backend:
cd backend
python -m venv venv

# Activate virtual environment
# Windows:
venv\Scripts\activate
# macOS/Linux:
source venv/bin/activate

# Install Python dependencies
pip install -r requirements.txt
  1. Initialize the database:
# The database will be automatically created on first run

Usage

Development Mode

  1. Start the Python backend:
cd backend
# Make sure virtual environment is activated
python main.py
  1. Start the Electron app in development mode:
# In the root directory
npm run dev

Production Build

  1. Build the application:
npm run build
  1. Package for distribution:
npm run package

Application Structure

handwriting-transcription-app/
├── src/                          # React frontend source
│   ├── components/               # UI components
│   ├── store/                   # State management
│   ├── types/                   # TypeScript definitions
│   └── main.tsx                 # Application entry point
├── electron/                    # Electron main process
│   ├── main.ts                  # Main process
│   └── preload.ts              # Preload script
├── backend/                     # Python backend
│   ├── main.py                  # FastAPI server
│   ├── services/               # Processing services
│   ├── models/                 # Data models
│   └── config/                 # Configuration
├── dist/                       # Build output
└── package.json               # Node.js configuration

Key Features Guide

File Processing

  1. Upload Files: Drag and drop or select PDF/image files
  2. Configure Options:
    • For PDFs: Specify page numbers (e.g., "1,3,5-10" or "all")
    • Enable/disable style learning
    • Adjust confidence threshold
  3. Process: Click "Process Files" to start recognition

Results Management

  • View Results: Access transcribed text in the Results panel
  • Edit Text: Click the edit icon to correct transcriptions
  • Export: Download results in various formats (TXT, JSON, CSV)
  • Form Elements: View detected checkboxes and radio buttons

Training Mode

  1. Add Samples: Upload handwriting images with correct transcriptions
  2. Validate Samples: Review and approve training data
  3. Train Model: Use validated samples to improve recognition accuracy
  4. Monitor Progress: Track model performance and accuracy metrics

Settings

  • Processing: Configure OCR engine, confidence thresholds, and concurrent jobs
  • Storage: Set default export formats and auto-save options
  • Privacy: Review data handling and clear training data if needed

Privacy & Security

  • Fully Local: All processing happens on your device
  • No Cloud Uploads: Documents never leave your computer
  • Encrypted Storage: Training data is securely stored locally
  • Automatic Cleanup: Temporary files are automatically removed

Troubleshooting

Common Issues

  1. Tesseract not found:

    • Ensure Tesseract is installed and in your system PATH
    • On Windows, you may need to set the path manually in settings
  2. Backend connection failed:

    • Check if Python backend is running on port 8000
    • Verify all Python dependencies are installed
  3. Poor OCR accuracy:

    • Ensure images are high quality and well-lit
    • Add training samples to improve recognition
    • Adjust preprocessing settings
  4. Performance issues:

    • Reduce concurrent job count in settings
    • Close other resource-intensive applications
    • Ensure sufficient RAM and disk space

System Requirements

  • Memory: 4GB RAM minimum, 8GB recommended
  • Storage: 2GB free disk space
  • CPU: Modern multi-core processor recommended
  • OS: Windows 10/11, macOS 10.14+, or Linux

Contributing

This is a local application designed for privacy and offline use. If you encounter issues or have suggestions:

  1. Check the troubleshooting section
  2. Review application logs in the developer console
  3. Ensure all dependencies are properly installed

License

This project is licensed under the MIT License - see the LICENSE file for details.

Acknowledgments

  • Tesseract OCR team for the powerful text recognition engine
  • OpenCV community for image processing capabilities
  • Electron team for the cross-platform desktop framework
  • All open-source contributors who made this project possible

About

Handwritten text transcription from PDF documents

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages