Metadata-Version: 2.4
Name: accsify-tesseract
Version: 1.0.5.1
Summary: Official Python SDK and CLI for the Accsify Monolithic Tesseract OCR Engine
Home-page: https://github.com/accsify/tesseract
Author: accsify
Author-email: accsify <contact@accsify.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/accsify/tesseract
Project-URL: Documentation, https://github.com/accsify/tesseract#readme
Project-URL: Repository, https://github.com/accsify/tesseract.git
Project-URL: Issues, https://github.com/accsify/tesseract/issues
Keywords: ocr,tesseract,text-recognition,computer-vision,accsify,layout-analysis
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: 3.15
Classifier: Topic :: Scientific/Engineering :: Image Recognition
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pillow>=9.0.0
Dynamic: author
Dynamic: home-page
Dynamic: license-file
Dynamic: requires-python

# accsify-tesseract

Official high-performance Python package and CLI for the **Accsify Tesseract OCR Engine**.

[![PyPI version](https://img.shields.io/badge/pypi-v5.5.0.1-blue.svg)](https://pypi.org/project/accsify-tesseract/)
[![Python Versions](https://img.shields.io/badge/python-3.8%20%7C%203.9%20%7C%203.10%20%7C%203.11%20%7C%203.12%20%7C%203.13%20%7C%203.14-blue.svg)](https://pypi.org/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Platform](https://img.shields.io/badge/platform-Windows%20x64%20%7C%20x86-green.svg)](https://microsoft.com)

---

## 1. Overview

`accsify-tesseract` is a modern, object-oriented Python SDK built on top of the monolithic `tesseract_engine.dll` Windows dynamic link library. It provides thread-safe OCR recognition, structured document layout hierarchy, native Windows WinHTTP streaming model downloads, multi-image batch processing, and multi-language support.

### Key Highlights
* **Zero Runtime External Dependencies**: The underlying native engine is statically linked with `/MT` (MSVC C/C++ Static Runtime). No Visual C++ Redistributable or external DLLs are needed.
* **Structured JSON & Dictionary Outputs**: Extract full page layout, words, confidences, bounding boxes, textline orders, and writing directions (LTR, RTL, TTB) directly into typed Python dictionaries or JSON.
* **Concurrent Model Flavors**: Download and store both **fast** (`tessdata_fast`) and **best** (`tessdata_best`) models simultaneously in separated subdirectories without file collisions.
* **Multi-Language OCR**: Seamlessly combine languages with `+` (e.g. `ara+eng`), with automatic downloading of missing language components.
* **Multi-Image Batch Processing**: High-throughput processing of image lists with unified structured results via `recognize_batch()`.
* **Native WinHTTP Downloader**: Download traineddata models with real-time progress callbacks and cancel tokens directly from official repositories over HTTPS.
* **Universal Image Support**: Directly load file paths (`str`, `Path`), raw encoded bytes (`bytes`, `bytearray`), PIL Images (`PIL.Image`), and NumPy image arrays (OpenCV `ndarray`).

---

## 2. Installation

### From PyPI
```bash
pip install accsify-tesseract
```

### From Local Source (Editable Mode)
```bash
cd d:\projects\c++\tesseract\python
pip install -e .
```

The package automatically discovers and loads `tesseract_engine.dll` from `dist/x64/`, `dist/x86/`, `bin/`, or standard Windows search paths based on whether you are running a 64-bit or 32-bit Python interpreter.

---

## 3. Quickstart

### One-Liner Recognition
```python
import accsify_tesseract as tess

# Extract plain text
text = tess.image_to_string("invoice.png", lang="eng")
print(text)

# Extract structured JSON string
json_str = tess.image_to_json("document.png", lang="eng")

# Extract structured dictionary with words and bounding boxes
data = tess.image_to_dict("document.png", lang="eng")
print(f"Mean confidence: {data['mean_confidence']}%")
for word in data["words"]:
    print(f"Word: {word['text']}, Box: {word['bbox']}, Conf: {word['confidence']}%")
```

---

## 4. Object-Oriented Engine Interface

The primary interface is `TesseractEngine`, implemented as a thread-safe context manager.

```python
from pathlib import Path
from accsify_tesseract import TesseractEngine, PageSegMode, ModelType

# Open engine with specific language and model flavor
with TesseractEngine(language="eng", flavor=ModelType.FAST) as engine:
    engine.set_page_seg_mode(PageSegMode.AUTO)
    
    # Load image from file path, bytes, PIL Image, or OpenCV ndarray
    engine.set_image(Path("document.png"))
    
    # Run recognition pipeline
    engine.recognize()
    
    # Retrieve outputs
    text = engine.get_text()
    mean_conf = engine.get_mean_confidence()
    hocr = engine.get_hocr(page_num=0)
    tsv = engine.get_tsv(page_num=0)
    
    print(f"Confidence: {mean_conf}%")
    print(text)
```

---

## 5. Structured JSON & Dictionary Extraction

`TesseractEngine.get_json()` and `TesseractEngine.get_structured_dict()` return deep structured layout information natively generated by the C++ engine:

```python
with TesseractEngine(language="eng") as engine:
    engine.set_image("sample.png")
    
    # Get Python dictionary
    doc = engine.get_structured_dict()
    
    print(f"Total Text: {doc['text']}")
    print(f"Mean Confidence: {doc['mean_confidence']}%")
    print(f"PSM: {doc['psm']}")
    
    for w in doc["words"]:
        print(f"Word: {w['text']:<20} Conf: {w['confidence']:<6.1f} Box: {w['bbox']} Direction: {w['direction']}")
```

### JSON Schema Specification
```json
{
  "text": "Full extracted UTF-8 document text...",
  "mean_confidence": 94,
  "psm": 3,
  "words": [
    {
      "text": "Accsify",
      "confidence": 98.4,
      "bbox": [120, 45, 230, 78],
      "direction": "LeftToRight",
      "order": "LTR",
      "deskew_angle": 0.00012
    }
  ]
}
```

---

## 6. Multi-Image Batch Processing

Process dozens or hundreds of images in a single session without recreating the engine:

```python
from accsify_tesseract import TesseractEngine

images = ["page1.png", "page2.png", "page3.png"]

with TesseractEngine(language="eng") as engine:
    batch_results = engine.recognize_batch(images)
    
    for res in batch_results:
        print(f"\n--- Result for: {res.image_path} ---")
        print(f"Confidence: {res.mean_confidence}% | Words: {res.words_count}")
        print(res.text[:150])  # preview first 150 chars
        
        # Export individual image result to JSON or dict
        json_payload = res.to_json(indent=2)
```

---

## 7. Multi-Language Combination & Auto-Download

You can combine multiple languages using the `+` operator. With `auto_download=True`, the engine automatically verifies and downloads any missing language before initializing:

```python
from accsify_tesseract import TesseractEngine, ModelType

# Combined Arabic and English recognition with auto-downloading
with TesseractEngine(language="ara+eng", flavor=ModelType.BEST, auto_download=True) as engine:
    engine.set_image("bilingual_contract.png")
    engine.recognize()
    print(engine.get_text())
```

---

## 8. Model Management & Storage Organization

Accsify Tesseract organizes models into structured subdirectories under the active `tessdata/` path to prevent collisions between different versions of the same language:

```
tessdata/
├── fast/                 # Compact models (~1-5 MB)
│   ├── eng.traineddata
│   └── ara.traineddata
├── best/                 # High-accuracy full LSTM models (~15-40 MB)
│   ├── eng.traineddata
│   └── ara.traineddata
├── script/               # Script-specific models
│   └── Arabic.traineddata
└── standard/             # Standard release models
```

### Querying Catalog & Downloading
```python
from accsify_tesseract import ModelManager, ModelType

# 1. Inspect installed models
installed = ModelManager.list_installed()
print("Installed models:", installed)

# 2. Query available catalog models
catalog = ModelManager.list_catalog(ModelType.FAST)
for m in catalog:
    print(f"{m.name:<15} {m.display_name:<30} {m.file_size_mb:.1f} MB (Installed: {m.is_installed})")

# 3. Live progress callback
def my_progress(name, model_type, downloaded, total, pct, status):
    print(f"\rDownloading {name}: {pct:.1f}% ({downloaded / 1024 / 1024:.2f} MB)", end="")
    return True  # return False to cancel download

# 4. Download fast Arabic model
ModelManager.download("ara", model_type=ModelType.FAST, progress_callback=my_progress)

# 5. Download best Arabic model (stored in tessdata/best/ara.traineddata without overwriting fast!)
ModelManager.download("ara", model_type=ModelType.BEST, progress_callback=my_progress)

# 6. Download multiple models in one call
ModelManager.download_many(["eng", "fra", "deu"], model_type=ModelType.FAST)

# 7. Configure custom storage directory
ModelManager.set_path("D:/my_models/tessdata")
print("Active path:", ModelManager.get_path())
```

---

## 9. Layout Analysis & Bounding Boxes

Analyze the full page geometry across 5 hierarchy levels: `BLOCK`, `PARA`, `TEXTLINE`, `WORD`, `SYMBOL`.

```python
from accsify_tesseract import TesseractEngine, PageIteratorLevel, WritingDirection

with TesseractEngine(language="ara+eng") as engine:
    engine.set_image("mixed_layout.png")
    layout = engine.analyse_layout(level=PageIteratorLevel.WORD)
    
    print(f"Discovered {len(layout.words)} words, {len(layout.lines)} lines")
    
    for word in layout.words:
        b = word.bbox
        is_rtl = word.writing_direction == WritingDirection.RIGHT_TO_LEFT
        dir_label = "RTL (Arabic)" if is_rtl else "LTR (Latin)"
        
        print(f"[{b.left}, {b.top}, {b.right}, {b.bottom}] {word.text:<25} Conf: {word.confidence:.1f}% Dir: {dir_label}")
    
    # Export full layout to JSON
    json_layout = layout.to_json(indent=2)
```

---

## 10. Orientation and Script Detection (OSD)

Detect page rotation degrees (0°, 90°, 180°, 270°) and identify the primary script:

```python
from accsify_tesseract import TesseractEngine, PageSegMode

with TesseractEngine(language="osd") as engine:
    engine.set_page_seg_mode(PageSegMode.OSD_ONLY)
    engine.set_image("scanned_page.png")
    
    osd = engine.detect_orientation_and_script()
    print(f"Orientation : {osd.orientation_deg}° (Confidence: {osd.orientation_confidence})")
    print(f"Script      : {osd.script_name} (Confidence: {osd.script_confidence})")
    print(f"Is Upright  : {osd.is_upright}")
    
    # Export to dict or JSON
    print(osd.to_dict())
```

---

## 11. Command-Line Interface (CLI)

The package provides the `accsify-tesseract` console script:

```bash
# 1. OCR a single image
accsify-tesseract ocr invoice.png -l eng

# 2. Batch OCR multiple images to structured JSON
accsify-tesseract ocr page1.png page2.png page3.png -l eng --format json -o batch_result.json

# 3. Multi-language OCR with high-accuracy 'best' model
accsify-tesseract ocr contract.png -l "ara+eng" --flavor best

# 4. Detailed layout inspection
accsify-tesseract layout document.png -l eng --format json

# 5. Detect page orientation and script
accsify-tesseract osd scanned.png

# 6. List and download models
accsify-tesseract models list --type fast
accsify-tesseract models download ara --type best
accsify-tesseract models installed
accsify-tesseract models path --set-path "D:\tessdata"
```

---

## 12. Exception Hierarchy

```python
from accsify_tesseract import (
    TesseractError,         # Base exception for all errors
    EngineInitError,        # Failed to initialize engine (missing traineddata, invalid path)
    ImageLoadError,         # Image file missing or corrupted
    RecognitionError,       # OCR recognition failure
    ModelDownloadError,     # Network / WinHTTP download failure
    ModelNotFoundError      # Requested model not found in catalog
)

try:
    with TesseractEngine(language="nonexistent") as engine:
        engine.set_image("sample.png")
except EngineInitError as e:
    print(f"Initialization error for {e.language} in {e.datapath}: {e}")
```

---

## 13. License

Distributed under the **MIT License**. Copyright (C) 2026 **accsify**. All rights reserved.
