TranscribeCppSharp

.NET bindings for transcribe.cpp: load GGUF speech-to-text models and transcribe audio (16 kHz mono float PCM) from C#.

This project is a packaging and binding effort only — the native library is developed by the transcribe.cpp authors and is not my work. See Attribution.

Two ways to use it:

  • A command-line tool — dotnet tool install -g TranscribeCppSharp.Cli gives you transcribe, which downloads a curated model on first use and prints or exports the transcript. See Command-line tool.
  • A .NET library — the TranscribeCppSharp wrapper (plus the native runtime package for your platform) for C# code. See Getting started.

Install

dotnet add package TranscribeCppSharp
dotnet add package TranscribeCppSharp.Native.linux-x64   # pick your platform

Or one package that works everywhere: dotnet add package TranscribeCppSharp.Bundle. The full platform list is under Installation.

The short version

using var model = Model.Load("model.gguf");
using var session = model.CreateSession();
var pcm = PcmExtensions.ReadWavToPcm("audio.wav");
var transcript = session.Run(pcm);
Console.WriteLine(transcript.FullText);

Model.Load picks the GPU when one initializes. The longer, explicit and deterministic form is on Getting started.

Try it without writing code

dotnet tool install -g TranscribeCppSharp.Cli
transcribe jfk.wav --model whisper-tiny

The model is fetched from HuggingFace on first use (42 MB for that alias), verified by sha256 and cached; later runs are offline. It runs on the GPU by default.

What it does not do

Stated plainly, so nothing is implied. Each item has its own page with the detail.

  • The library is not thread-safe. At most one run per model at a time, across all of that model’s sessions; parallel workers each need their own Model. This is an upstream 0.x limitation, not one this wrapper imposes (Concurrency).
  • Streaming cannot attribute speakers. transcribe_stream_params has no diarize field in transcribe.cpp v0.2.4, so the streaming API cannot request it; we do not invent one (Diarization).
  • No CUDA runtime ships in the packages. The bundled binaries are CPU + Vulkan on Windows/Linux and Metal on macOS; an NVIDIA GPU means supplying your own CUDA build (Using CUDA).
  • The MIT license does not cover the models. They come from different ecosystems, some non-commercial, and this project neither bundles nor curates them (Model licenses).
  • Speaker attribution is not verified in CI. It is covered only when the ~617 MB MOSS asset is fetched with WITH_DIARIZATION_MODEL=1. Until then, a regression that emptied SpeakerSegments would go unnoticed (diarization).
  • There is no GPU in CI. The tool is tested end to end on the CPU, but GPU selection itself is covered by policy tests, not by a real run (what the CLI does not do).

Where to go next

Page What is in it
Getting started installing, the native packages, load-and-transcribe, batch, streaming, sample code
Command-line tool installing transcribe, a first run, the full --help, and what the tool does not do
Compute GPU by default, --list-devices, backend and device selection, CUDA
Models aliases vs. HuggingFace specs, --list-models, caching, capabilities, model licences
Exporting --out, and the plain, WebVTT and JSON formats
Long audio and diarization windowing, the four diarization entry points, and what is not verified
Audio input which formats are tested, ffmpeg, and the measured memory behaviour
Real-time streaming StreamSession, partial results, its two limits
Concurrency blocking calls, thread safety per type, dispose discipline
Architecture the layers, native library resolution, versioning vs. upstream
Error handling TranscribeException, StatusCode, the failures you will actually meet
Development building, testing, the checks that keep these docs honest, building from source
Governance security, licence, attribution

Source on GitHub · NuGet · CHANGELOG · MIT


This site uses Just the Docs, a documentation theme for Jekyll.