16 KiB
DocEngine — Document Template Recognition & Reconstruction System
A production-ready system for scanning documents, detecting layouts, extracting content, generating reusable templates, matching future uploads against stored templates, and reconstructing original layouts as PDF output.
Architecture
┌───────────────┐ ┌───────────────┐ ┌──────────────┐
│ FastAPI App │─────▶│ Celery │─────▶│ Redis │
│ (Port 7989) │ │ Worker(s) │ │ (Broker) │
└───────┬───────┘ └───────┬───────┘ └──────────────┘
│ │
▼ ▼
┌───────────────────────────────────────┐
│ PostgreSQL (Schema: admin) │
│ 192.168.0.111:5432 │
└───────────────────────────────────────┘
Stack: Python 3.12, FastAPI, SQLAlchemy 2.x, Pydantic V2, Celery, Redis, PaddleOCR, PyMuPDF, OpenCV, ReportLab, PostgreSQL 16.
Features
| Capability | Implementation |
|---|---|
| Scanned image OCR | PaddleOCR (CPU/GPU) |
| Native PDF text extraction | PyMuPDF (fitz) |
| Layout detection | OpenCV + LayoutParser |
| Table extraction | Camelot-py + OpenCV contour detection |
| Header/footer detection | Positional heuristics |
| Watermark detection | Transparency + large-font analysis |
| Font info extraction | PyMuPDF text dict parsing |
| Template generation & storage | PostgreSQL (admin schema) |
| Template fingerprinting | SHA-256 structural hashing |
| Template matching | Multi-signal similarity scoring |
| PDF reconstruction | ReportLab from template definitions |
| Async processing | Celery + Redis |
| Authentication | JWT (access + refresh tokens, bcrypt) |
| Monitoring | Prometheus + structlog JSON logging |
Project Structure
docengine/
├── app/
│ ├── main.py # FastAPI application entry
│ ├── api/
│ │ ├── router.py # Top-level API router
│ │ └── v1/
│ │ ├── auth.py # Auth endpoints
│ │ ├── documents.py # Document endpoints
│ │ ├── health.py # Health check
│ │ └── templates.py # Template endpoints
│ ├── core/
│ │ ├── config.py # Pydantic Settings
│ │ ├── database.py # SQLAlchemy engine & session
│ │ ├── dependencies.py # FastAPI DI
│ │ ├── exceptions.py # Custom exception hierarchy
│ │ ├── logging_config.py # structlog configuration
│ │ └── security.py # JWT & bcrypt helpers
│ ├── models/ # SQLAlchemy ORM models
│ ├── schemas/ # Pydantic request/response schemas
│ ├── repositories/ # Data access layer
│ ├── services/ # Business logic
│ │ ├── document_service.py # Orchestration pipeline
│ │ ├── ocr_service.py # PaddleOCR integration
│ │ ├── pdf_service.py # PyMuPDF native PDF parsing
│ │ ├── layout_service.py # OpenCV layout detection
│ │ ├── template_service.py # Template generation
│ │ ├── fingerprint_service.py
│ │ ├── matching_service.py
│ │ └── reconstruction_service.py
│ ├── middleware/ # CORS, audit, metrics, rate limit
│ ├── storage/ # File storage abstraction
│ ├── tasks/ # Celery async tasks
│ ├── workers/ # Celery app configuration
│ └── events/ # App lifecycle handlers
├── alembic/ # Database migrations
├── sql/ # Raw SQL scripts
├── tests/ # Test suite
├── docker-compose.yml # Dev stack
├── docker-compose.prod.yml # Production stack
├── Dockerfile # Multi-stage build
├── requirements.txt
└── .env
Quick Start
Prerequisites
- Python 3.12+
- PostgreSQL 16 (running at
192.168.0.111:5432) - Redis (for Celery)
poppler-utilsandghostscript(for pdf2image/camelot)
Local Setup
# Clone & enter
cd docengine
# Create virtual environment
python -m venv .venv
source .venv/bin/activate
# Install dependencies
pip install -r requirements.txt
# Create storage directories
mkdir -p storage/{documents,templates,images,temp,rendered}
# Run database migrations
alembic upgrade head
# (Optional) Seed default data
psql -h 192.168.0.111 -p 5432 -U postgres -d ocr -f sql/003_seed_data.sql
psql -h 192.168.0.111 -p 5432 -U postgres -d ocr -f sql/004_indexes.sql
# Start the application
python -m app.main
The API is now available at http://localhost:7989. Interactive docs at http://localhost:7989/docs.
Start Celery Worker (separate terminal)
source .venv/bin/activate
celery -A app.workers.celery_app worker --loglevel=info --concurrency=4
Docker Setup
# Build and start all services (app + worker + db + redis)
docker compose up --build -d
# Run migrations inside the container
docker compose exec app alembic upgrade head
# Seed data
docker compose exec app bash -c "psql -h db -U postgres -d ocr -f sql/003_seed_data.sql"
API Reference
Base URL: http://localhost:7989/api/v1
Health
curl http://localhost:7989/api/v1/health
Authentication
Register
curl -X POST http://localhost:7989/api/v1/auth/register \
-H "Content-Type: application/json" \
-d '{
"username": "john",
"email": "john@example.com",
"password": "SecurePass123!",
"full_name": "John Doe"
}'
Login
curl -X POST http://localhost:7989/api/v1/auth/login \
-H "Content-Type: application/json" \
-d '{
"username": "john",
"password": "SecurePass123!"
}'
Response:
{
"access_token": "eyJhbGciOiJIUzI1NiIs...",
"refresh_token": "eyJhbGciOiJIUzI1NiIs...",
"token_type": "bearer",
"expires_in": 1800
}
Get Current User
curl http://localhost:7989/api/v1/auth/me \
-H "Authorization: Bearer <access_token>"
Refresh Token
curl -X POST http://localhost:7989/api/v1/auth/refresh \
-H "Content-Type: application/json" \
-d '{"refresh_token": "<refresh_token>"}'
Change Password
curl -X POST http://localhost:7989/api/v1/auth/change-password \
-H "Authorization: Bearer <access_token>" \
-H "Content-Type: application/json" \
-d '{
"current_password": "SecurePass123!",
"new_password": "NewSecurePass456!"
}'
Logout
curl -X POST http://localhost:7989/api/v1/auth/logout \
-H "Authorization: Bearer <access_token>" \
-H "Content-Type: application/json" \
-d '{"refresh_token": "<refresh_token>"}'
Documents
Upload Document
# Upload a PDF
curl -X POST http://localhost:7989/api/v1/documents/upload \
-H "Authorization: Bearer <access_token>" \
-F "file=@/path/to/document.pdf"
# Upload a scanned image
curl -X POST http://localhost:7989/api/v1/documents/upload \
-H "Authorization: Bearer <access_token>" \
-F "file=@/path/to/scan.jpg"
Response:
{
"id": "550e8400-e29b-41d4-a716-446655440000",
"filename": "abc123_document.pdf",
"original_filename": "document.pdf",
"content_type": "application/pdf",
"file_size": 245760,
"checksum": "e3b0c44298fc1c149afbf4c8996fb924...",
"status": "pending",
"created_at": "2026-06-01T12:00:00Z"
}
Get Document
curl http://localhost:7989/api/v1/documents/<document_id> \
-H "Authorization: Bearer <access_token>"
List Documents
# With pagination
curl "http://localhost:7989/api/v1/documents?page=1&page_size=20" \
-H "Authorization: Bearer <access_token>"
# Filter by status
curl "http://localhost:7989/api/v1/documents?status=completed" \
-H "Authorization: Bearer <access_token>"
Delete Document
curl -X DELETE http://localhost:7989/api/v1/documents/<document_id> \
-H "Authorization: Bearer <access_token>"
Get Template Matches for Document
curl http://localhost:7989/api/v1/documents/<document_id>/template \
-H "Authorization: Bearer <access_token>"
Templates
List Templates
curl "http://localhost:7989/api/v1/templates?page=1&page_size=20" \
-H "Authorization: Bearer <access_token>"
Get Template
curl http://localhost:7989/api/v1/templates/<template_id> \
-H "Authorization: Bearer <access_token>"
Delete (Deactivate) Template
curl -X DELETE http://localhost:7989/api/v1/templates/<template_id> \
-H "Authorization: Bearer <access_token>"
Match Document to Templates
curl -X POST http://localhost:7989/api/v1/templates/match \
-H "Authorization: Bearer <access_token>" \
-H "Content-Type: application/json" \
-d '{
"document_id": "<document_id>",
"min_confidence": 0.5,
"max_results": 5
}'
Response:
[
{
"id": "...",
"document_id": "...",
"format_id": "...",
"confidence_score": 0.92,
"match_details": { "dimension_score": 1.0, "header_score": 0.85 },
"selected": true,
"template_name": "Invoice Template v1",
"created_at": "2026-06-01T12:00:00Z"
}
]
Render Template to PDF
curl -X POST http://localhost:7989/api/v1/templates/render \
-H "Authorization: Bearer <access_token>" \
-H "Content-Type: application/json" \
-d '{
"template_id": "<template_id>",
"data": {
"company_name": "Acme Corp",
"invoice_number": "INV-2026-001",
"date": "2026-06-01",
"total": "$1,250.00"
},
"output_filename": "invoice_output.pdf"
}'
Response:
{
"output_path": "rendered/invoice_output.pdf",
"filename": "invoice_output.pdf",
"file_size": 32768,
"page_count": 1,
"rendered_at": "2026-06-01T12:05:00Z"
}
Download Rendered PDF
curl -O http://localhost:7989/api/v1/templates/<template_id>/download?filename=invoice_output.pdf \
-H "Authorization: Bearer <access_token>"
Processing Pipeline
When a document is uploaded, the following Celery task pipeline executes asynchronously:
- File Type Detection — Determine if the document is a native PDF or scanned image.
- Page Extraction — Convert PDF pages to images (for scanned docs) or parse directly (for native PDFs).
- OCR — Run PaddleOCR on scanned pages to extract text blocks with coordinates, confidence, and bounding boxes.
- Native PDF Parsing — Use PyMuPDF to extract text, fonts, images, and tables from native PDFs.
- Layout Analysis — Detect headers, footers, watermarks, tables, and image regions using OpenCV heuristics.
- Template Generation — Build a reusable template definition from the detected layout, stored in PostgreSQL.
- Fingerprint Generation — Compute a structural fingerprint (SHA-256) for future matching.
- Status Update — Mark the document as
completed(orfailedwith error details).
Database
Connection: postgresql://postgres:***@192.168.0.111:5432/ocr
Schema: admin
Migrations
# Create a new migration
alembic revision --autogenerate -m "description"
# Apply migrations
alembic upgrade head
# Rollback one step
alembic downgrade -1
Tables
| Table | Purpose |
|---|---|
users |
User accounts |
roles |
Role definitions (admin, user, viewer) |
user_roles |
User-role mapping (M2M) |
refresh_tokens |
JWT refresh token storage |
audit_logs |
Action audit trail |
documents |
Uploaded document records |
document_pages |
Per-page data (dimensions, images) |
document_text_blocks |
Extracted text with position & font info |
document_images |
Extracted images with position |
document_tables |
Extracted tables with cell data (JSONB) |
document_formats |
Template definitions |
document_cells |
Template cell layout definitions |
document_regions |
Template region definitions |
table_formats |
Template table structure definitions |
table_columns |
Template table column definitions |
table_rows |
Template table row definitions |
watermarks |
Template watermark definitions |
image_regions |
Template image region definitions |
template_fingerprints |
Structural fingerprints for matching |
template_matches |
Document-to-template match results |
Testing
# Install dev dependencies
pip install -r requirements-dev.txt
# Run all tests
pytest
# Run with coverage
pytest --cov=app --cov-report=term-missing
# Run specific test categories
pytest tests/unit/
pytest tests/api/
pytest tests/repositories/
Configuration
All configuration is via environment variables (.env file). Key settings:
| Variable | Default | Description |
|---|---|---|
APP_PORT |
7989 |
Application port |
DB_HOST |
192.168.0.111 |
PostgreSQL host |
DB_PORT |
5432 |
PostgreSQL port |
DB_NAME |
ocr |
Database name |
DB_SCHEMA |
admin |
PostgreSQL schema |
REDIS_HOST |
192.168.0.111 |
Redis host |
CELERY_BROKER_URL |
redis://:***@192.168.0.111:7901/0 |
Celery broker |
JWT_SECRET_KEY |
(see .env) | JWT signing key |
JWT_ACCESS_TOKEN_EXPIRE_MINUTES |
30 |
Access token TTL |
STORAGE_LOCAL_PATH |
./storage |
Local file storage path |
STORAGE_MAX_FILE_SIZE_MB |
100 |
Max upload size |
OCR_LANGUAGE |
en |
PaddleOCR language |
OCR_USE_GPU |
false |
Enable GPU for OCR |
Production Deployment
# Using production compose file
docker compose -f docker-compose.prod.yml up --build -d
# Scale workers
docker compose -f docker-compose.prod.yml up --scale worker=4 -d
Production compose includes:
- Resource limits (CPU/memory)
- Redis authentication
- App replicas
- Persistent named volumes
- Auto-restart policies
Default Credentials
| Username | Password | Role |
|---|---|---|
admin |
Admin@123! |
admin |
⚠️ Change the default admin password immediately in production.
License
Proprietary — All rights reserved.