Metadata-Version: 2.4
Name: accsify-whisper
Version: 1.0.0.2
Summary: High-performance Whisper Speech Recognition Engine for Windows with C++ acceleration
Home-page: https://accsify.com
Author: accsify
Author-email: accsify <support@accsify.com>
License-Expression: MIT
Project-URL: Homepage, https://accsify.com
Project-URL: Repository, https://github.com/accsify/whisper
Keywords: whisper,speech-to-text,transcription,audio,voice,ai,speech-recognition,windows,whisper.cpp,accsify
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Developers
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: 3.15
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.8
Description-Content-Type: text/markdown
Dynamic: author
Dynamic: home-page
Dynamic: requires-python

# Accsify Whisper Engine for Python (`accsify_whisper`)

Enterprise-grade Windows speech recognition SDK powered by **`whisper.cpp`** and **Windows Media Foundation**, developed and branded by **accsify**.

This library provides both an object-oriented and a high-level functional Python interface to the native `whisper_engine.dll`, delivering high-speed offline inference, dual architecture auto-detection (x64 / x86), universal zero-dependency audio decoding, automatic hardware diagnostics, background WinHTTP model downloading, cloud API rotation, and real-time streaming pipelines.

---

## Table of Contents

- [Key Highlights](#key-highlights)
- [Architecture & Dual-Binary Auto-Detection](#architecture--dual-binary-auto-detection)
- [Installation](#installation)
- [Quick Start: Functional vs OOP API](#quick-start-functional-vs-oop-api)
- [Feature Breakdown & Code Walkthroughs](#feature-breakdown--code-walkthroughs)
  - [1. Hardware Inspection & SIMD Diagnostics](#1-hardware-inspection--simd-diagnostics)
  - [2. Model Catalog, Directory Scanner & WinHTTP Downloader](#2-model-catalog-directory-scanner--winhttp-downloader)
  - [3. Universal Audio Decoding & Format Conversion](#3-universal-audio-decoding--format-conversion)
  - [4. Local Offline Model Transcription](#4-local-offline-model-transcription)
  - [5. Advanced Transcription (VAD, DTW Word Timestamps, Prompting)](#5-advanced-transcription-vad-dtw-word-timestamps-prompting)
  - [6. Subtitle & Multi-Format Exports (SRT, VTT, JSON)](#6-subtitle--multi-format-exports-srt-vtt-json)
  - [7. Real-Time Audio Streaming Pipeline](#7-real-time-audio-streaming-pipeline)
  - [8. Online Cloud Whisper API with Multi-Key Rotation](#8-online-cloud-whisper-api-with-multi-key-rotation)
  - [9. Hybrid Transcription (Local-First with Cloud Failover)](#9-hybrid-transcription-local-first-with-cloud-failover)
  - [10. Concurrency, Multi-Threading & Context Managers](#10-concurrency-multi-threading--context-managers)
- [Bundled Test & Example Suite (`accsify_whisper.tests`)](#bundled-test--example-suite-accsify_whispertests)
- [Exception Hierarchy](#exception-hierarchy)
- [PyPI Packaging & Distribution](#pypi-packaging--distribution)
- [License & Company Information](#license--company-information)

---

## Key Highlights

- **Dual-Architecture Native Binaries Bundled**: Includes pre-compiled `whisper_engine.dll` for both **64-bit (x64)** and **32-bit (x86)** Windows environments. No C++ compiler or CMake required on client systems.
- **Zero-Configuration Bitness Auto-Detection**: Automatically queries interpreter pointer size at runtime (`struct.calcsize("P") == 8`) and loads the exact matching DLL. Eliminates WinError 193 mismatches completely.
- **Zero-Dependency Audio Decoding**: Native Windows Media Foundation (WMF) decodes **MP3, M4A, AAC, WAV, WMA, and FLAC** directly into 16,000 Hz 32-bit float mono PCM. No external `ffmpeg.exe` binary needed!
- **Automatic SIMD Exploitation**: Native CPUID inspection detects AVX, AVX2, AVX-512, FMA, and F16C CPU vector extensions.
- **Full Model Registry (34 Models)**: Complete catalog from `tiny` to `large-v3-turbo`, including 5-bit quantized models (**Q5_0, Q5_1, Q8_0**) and English-only models.
- **Asynchronous WinHTTP Downloader**: Non-blocking downloads with live speed (KB/s), progress percentage, and ETA calculation.
- **High-Precision Word Timestamps (DTW)**: Dynamic Time Warping token alignment for millisecond-level word timing.
- **Subtitle Generation**: Native export to SubRip (`.srt`) and WebVTT (`.vtt`) formats.
- **Built-in Standalone Test & Example Suite**: Includes 10 dedicated test files with comprehensive real-world scenarios installed directly with the package.

---

## Architecture & Dual-Binary Auto-Detection

The package bundles native binaries in architecture-specific folders:
```text
accsify_whisper/
├── libraries/
│   ├── x64/
│   │   └── whisper_engine.dll    (64-bit native core)
│   └── x86/
│       └── whisper_engine.dll    (32-bit native core)
```

When initializing `WhisperEngine()`, the loader automatically resolves the correct DLL:
1. Custom `dll_path` parameter (if passed).
2. `ACCSIFY_WHISPER_DLL` environment variable.
3. `accsify_whisper/libraries/<arch>/whisper_engine.dll` (matching 64-bit or 32-bit Python).
4. Fallback search paths in `build/dist/<arch>/` and `build/bin/`.

---

## Installation

### From PyPI
```bash
pip install accsify-whisper
```

### From Local Source / Wheel
```bash
cd wrappers/python
pip install .
```

---

## Quick Start: Functional vs OOP API

`accsify_whisper` provides both high-level functional shortcuts and an OOP engine:

### 1. Functional Shortcuts
```python
import accsify_whisper as aw

# Hardware check
hw = aw.check_system_support()
print(f"Supported: {hw.supported}, CPU: {hw.cpu_brand}, Threads: {hw.recommended_threads}")

# List models
models = aw.list_available_models()
print(f"Available models: {len(models)}")

# Decode any audio file (MP3, M4A, WAV, etc.) to 16kHz float samples
samples = aw.decode_audio_file("recording.mp3")

# Convert audio to normalized 16kHz WAV
aw.convert_audio("speech.m4a", "speech_16k.wav")
```

### 2. Object-Oriented Engine
```python
from accsify_whisper import WhisperEngine

# Initialize engine (auto-detects x64/x86 DLL)
engine = WhisperEngine()

# Download tiny model with live progress
model_path = engine.models.download("tiny-q5_1", destination_folder="./models")

# Transcribe with context manager
with engine.load_model(model_path) as model:
    result = model.transcribe_file("recording.mp3", language="en")
    print(f"Transcript: {result.text}")
    
    # Save subtitles
    result.save_srt("subtitles.srt")
    result.save_vtt("subtitles.vtt")
```

---

## Feature Breakdown & Code Walkthroughs

### 1. Hardware Inspection & SIMD Diagnostics

```python
from accsify_whisper import check_system_support, get_system_info_json

# Structured dataclass
hw = check_system_support()
print(f"Supported on PC:     {hw.is_supported}")
print(f"CPU Brand:           {hw.cpu_brand}")
print(f"Physical/Logical:    {hw.physical_cores} cores / {hw.logical_cores} threads")
print(f"Recommended Threads: {hw.recommended_threads}")
print(f"SIMD:                AVX={hw.has_avx}, AVX2={hw.has_avx2}, FMA={hw.has_fma}")
print(f"Physical RAM:        {hw.total_ram_mb:,} MB ({hw.available_ram_mb:,} MB free)")

for idx, gpu in enumerate(hw.gpus, 1):
    print(f"GPU #{idx}: {gpu.name} (Dedicated VRAM: {gpu.dedicated_vram_mb:,} MB)")

print(f"Recommended Models:  {hw.recommended_models}")

# Raw JSON report
json_report = get_system_info_json()
```

---

### 2. Model Catalog, Directory Scanner & WinHTTP Downloader

```python
from accsify_whisper import WhisperEngine

engine = WhisperEngine()

# 1. Query full 34-model catalog
models = engine.models.list_available(local_cache_dir="./models")
for m in models[:5]:
    print(f"{m.id:<18} | Size: {m.disk_size_mb:>4} MB | Quant: {m.quantization} | Downloaded: {m.is_downloaded}")

# 2. Lookup model by key
info = engine.models.get("tiny-q5_1")
print(f"Model URL: {info.download_url}")

# 3. Scan local directory for downloaded models
found = engine.models.scan_directory("./models")
for lm in found:
    print(f"Found: {lm.filename} ({lm.size_mb:.1f} MB, valid={lm.is_valid}, format={lm.format})")

# 4. Verify binary magic bytes (GGML 0x67676d6c / GGUF 0x47475546)
is_valid = engine.models.verify_file("./models/ggml-tiny-q5_1.bin")

# 5. Download model with live progress
def on_progress(p):
    print(f"\rDownloading: {p.percent:.1f}% ({p.speed_kbps:.1f} KB/s, ETA: {p.eta_seconds}s)", end="")

path = engine.models.download("tiny-q5_1", destination_folder="./models", on_progress=on_progress)
```

---

### 3. Universal Audio Decoding & Format Conversion

Zero external dependencies! Native Windows Media Foundation integration handles **MP3, M4A, AAC, WAV, WMA, and FLAC**:

```python
from accsify_whisper import decode_audio_file, decode_audio_buffer, convert_audio

# Decode audio file to 16,000 Hz 32-bit float mono PCM samples
samples = decode_audio_file("meeting.mp3")
print(f"Decoded {len(samples)} samples ({len(samples)/16000:.2f} seconds)")

# Decode from in-memory bytes
with open("voice_note.m4a", "rb") as f:
    raw_bytes = f.read()
mem_samples = decode_audio_buffer(raw_bytes)

# Convert any audio file to normalized 16kHz WAV on disk
convert_audio("speech.aac", "speech_16k.wav")
```

---

### 4. Local Offline Model Transcription & Smart Model Auto-Selection

```python
from accsify_whisper import WhisperEngine, get_best_available_model

engine = WhisperEngine()

# Smart Model Auto-Selection:
# Omit model_path or pass None/"" to automatically scan and pick the highest-tier model
# matching your system's available RAM from ./models, ./, or executable paths!
best_model_path = engine.get_best_model()  # or get_best_available_model()
print(f"Auto-selected best model: {best_model_path}")

# Load model (explicit path or automatic selection)
with engine.load_model() as model:  # auto-detects best available local model
    # Transcribe file
    result = model.transcribe_file(
        "audio.wav",
        language="auto",        # "en", "es", "fr", "ar", "auto"
        translate=False,        # Translate foreign speech to English
        n_threads=4,            # CPU inference threads (0 = auto-detect optimal)
        strategy="greedy",      # "greedy" or "beam_search"
        beam_size=5,            # Beam search width
        temperature=0.0,
        temperature_inc=0.2,
    )
    print(f"Text: {result.text}")
    print(f"Detected Language: {result.language}")

    # Access segmented timestamps
    for seg in result.segments:
        print(f"[{seg.start_seconds:.2f}s -> {seg.end_seconds:.2f}s] {seg.text}")

    # Transcribe raw float samples directly from memory
    result_pcm = model.transcribe_pcm(samples, language="auto")
```

---

### 5. Advanced Transcription (VAD, DTW Word Timestamps, Prompting)

```python
with engine.load_model("./models/ggml-tiny-q5_1.bin", dtw_token_timestamps=True) as model:
    # Real-time callbacks
    def on_segment(seg_data):
        print(f"  [Partial]: {seg_data.get('text', '')}")

    def on_progress(pct):
        print(f"Progress: {pct}%")

    result = model.transcribe_file(
        "interview.mp3",
        word_timestamps=True,       # Exact word-level DTW timestamps
        vad_enable=False,           # Voice activity detection
        initial_prompt="Acme Corp quarterly earnings report, CEO John Doe", # Domain priming
        on_progress=on_progress,
        on_segment=on_segment,
    )

    # Word-level breakdown
    for seg in result.segments:
        print(f"[{seg.start_ms}ms -> {seg.end_ms}ms] {seg.text}")
        for w in seg.words:
            print(f"    - '{w.word}' ({w.start_ms}ms - {w.end_ms}ms, prob: {w.probability:.2f})")
```

---

### 6. Subtitle & Multi-Format Exports (SRT, VTT, JSON)

```python
# 1. SubRip Subtitles (.srt)
srt_str = result.to_srt()
result.save_srt("output.srt")

# 2. WebVTT Subtitles (.vtt)
vtt_str = result.to_vtt()
result.save_vtt("output.vtt")

# 3. Structured JSON Schema
json_str = result.to_json(parse=False)
result.save_json("output.json")

# 4. Parsed JSON dictionary with timings
data = result.to_json(parse=True)
print("Encoder latency:", result.timings.get("encode_ms"), "ms")
print("Total latency:  ", result.timings.get("total_ms"), "ms")
```

---

### 7. Real-Time Audio Streaming Pipeline

Process streaming microphone chunks with sliding window buffer:

```python
with engine.load_model("./models/ggml-tiny-q5_1.bin") as model:
    # 1s step size, 4s rolling window
    stream = model.create_stream(step_ms=1000, length_ms=4000, language="en")
    
    # Push chunks as they arrive from mic (e.g. 500ms chunks = 8000 samples)
    for chunk in audio_stream_generator():
        stream.feed(chunk)
        new_text = stream.process()
        if new_text:
            print(f"Live: {new_text}", end="\r")
            
    # Retrieve final merged transcription
    full_transcript = stream.get_full_text()
    stream.close()
```

---

### 8. Online Cloud Whisper API with Multi-Key Rotation

Execute cloud transcriptions with automatic key rotation and failover on HTTP 429 rate limits:

```python
from accsify_whisper.cloud import CloudConfig

cloud_cfg = CloudConfig(
    endpoint_url="https://api.openai.com/v1/audio/transcriptions",
    api_keys=["sk-primary-key...", "sk-backup-key-1...", "sk-backup-key-2..."],
    model_name="whisper-1",
    language="en",
    timeout_seconds=30
)

# Transcribe via cloud
cloud_result = engine.cloud.transcribe_file("audio.mp3", config=cloud_cfg)
print("Cloud Transcript:", cloud_result.text)
```

---

### 9. Hybrid Transcription (Local-First with Cloud Failover)

Attempt local transcription on host hardware first; if local hardware is overloaded or lacks models, seamlessly fall back to cloud API:

```python
with engine.load_model("./models/ggml-tiny-q5_1.bin") as model:
    result = engine.cloud.transcribe_hybrid(
        audio_filepath="audio.mp3",
        config=cloud_cfg,
        model=model,
        language="en"
    )
    print("Hybrid Transcript:", result.text)
```

---

### 10. Concurrency, Multi-Threading & Context Managers

```python
# 1. Multiple concurrent engine instances
engine_a = WhisperEngine()
engine_b = WhisperEngine()

# 2. Context manager ensures native C++ memory is freed immediately
with engine.load_model("./models/ggml-tiny-q5_1.bin") as model:
    res = model.transcribe_file("sample.wav")

# 3. Dynamic model swapping on a single engine
model1 = engine.load_model("./models/ggml-tiny-q5_1.bin")
# ... use model1 ...
model1.close()

model2 = engine.load_model("./models/ggml-base-q5_1.bin")
# ... use model2 ...
model2.close()
```

---

## Bundled Test & Example Suite (`accsify_whisper.tests`)

`accsify_whisper` comes with a comprehensive, standalone test and example suite installed directly with the package under `accsify_whisper.tests`:

| Module | Feature Tested | Scenarios Covered |
|:---|:---|:---|
| [`test_01_hardware_inspection.py`](accsify_whisper/tests/test_01_hardware_inspection.py) | Hardware Diagnostics | PC compatibility, CPU SIMD flags, RAM metrics, GPU VRAM detection, recommended models, JSON report |
| [`test_02_model_catalog.py`](accsify_whisper/tests/test_02_model_catalog.py) | Model Catalog | 34-model registry, lookup by key, 5-bit Q5 filter, English-only filter, folder scan, binary validation |
| [`test_03_model_download.py`](accsify_whisper/tests/test_03_model_download.py) | WinHTTP Downloader | Live console progress callback, custom folder download, caching check, cancel/abort simulation |
| [`test_04_transcription_basic.py`](accsify_whisper/tests/test_04_transcription_basic.py) | Basic Transcription | WAV file, memory PCM float samples, language selection, auto-detect, translation, thread scaling |
| [`test_05_transcription_advanced.py`](accsify_whisper/tests/test_05_transcription_advanced.py) | Advanced Options | VAD enable/disable, DTW word timestamps, initial prompt, greedy vs beam search, temperature, callbacks |
| [`test_06_audio_decoding.py`](accsify_whisper/tests/test_06_audio_decoding.py) | Audio Decoding (WMF) | 16-bit PCM WAV, 32-bit float WAV, MP3 decoding, in-memory bytes, format conversion, corrupt buffer handling |
| [`test_07_subtitles_and_exports.py`](accsify_whisper/tests/test_07_subtitles_and_exports.py) | Subtitles & Exports | SRT export, WebVTT export, structured JSON schema, encoder/decoder timings breakdown, plain text |
| [`test_08_streaming.py`](accsify_whisper/tests/test_08_streaming.py) | Real-Time Streaming | Stream session initialization, simulated chunk feeding, live transcript updates, final text merging |
| [`test_09_cloud_fallback.py`](accsify_whisper/tests/test_09_cloud_fallback.py) | Cloud & Hybrid Client | CloudConfig setup, multi-key rotation, CloudWhisperClient execution, hybrid local-first fallback |
| [`test_10_concurrency_and_lifecycle.py`](accsify_whisper/tests/test_10_concurrency_and_lifecycle.py) | Concurrency & Lifecycle | Multiple engines, multi-threaded parallel workers, model swapping, context manager, close idempotency |
| [`test_all.py`](accsify_whisper/tests/test_all.py) | Master Runner | Runs all 10 feature test suites sequentially with formatted summary table |

### Running the Tests / Examples

Run the entire master test suite:
```bash
python -m accsify_whisper.tests.test_all
```

Or run any individual feature example standalone:
```bash
python -m accsify_whisper.tests.test_01_hardware_inspection
python -m accsify_whisper.tests.test_04_transcription_basic
python -m accsify_whisper.tests.test_06_audio_decoding
python -m accsify_whisper.tests.test_08_streaming
```

---

## Exception Hierarchy

All exceptions inherit from `accsify_whisper.exceptions.WhisperError`:

```text
WhisperError
├── HardwareNotSupportedError     # Host PC lacks required SIMD instruction sets
├── ModelLoadError                # Failure loading or parsing model weights
├── AudioDecodeError              # Media Foundation failed to decode audio file/buffer
├── InferenceError                # whisper.cpp transcription failure
├── DownloadError                 # WinHTTP network or disk write failure
└── CloudApiError                 # Cloud API error after all keys exhausted
```

---


## License & Company Information

- **Brand / Company**: accsify
- **Product**: Accsify Whisper Engine Python SDK
- **Version**: 1.0.0.1
- **Platform**: Windows (x64, x86)
- **License**: MIT License
- **Support**: support@accsify.com
- **Website**: [https://accsify.com](https://accsify.com)
