TranscribeCppSharp
.NET bindings for transcribe.cpp: load GGUF speech-to-text models and transcribe audio (16 kHz mono float PCM) from C#.
This project is a packaging and binding effort only — the native library is developed by the transcribe.cpp authors and is not my work. See Attribution.
Two ways to use it:
- A command-line tool —
dotnet tool install -g TranscribeCppSharp.Cligives youtranscribe, which downloads a curated model on first use and prints or exports the transcript. See Command-line tool. - A .NET library — the
TranscribeCppSharpwrapper (plus the native runtime package for your platform) for C# code. See Getting started.
Install
dotnet add package TranscribeCppSharp
dotnet add package TranscribeCppSharp.Native.linux-x64 # pick your platform
Or one package that works everywhere: dotnet add package TranscribeCppSharp.Bundle. The full platform list is under Installation.
The short version
using var model = Model.Load("model.gguf");
using var session = model.CreateSession();
var pcm = PcmExtensions.ReadWavToPcm("audio.wav");
var transcript = session.Run(pcm);
Console.WriteLine(transcript.FullText);
Model.Load picks the GPU when one initializes. The longer, explicit and deterministic form is on Getting started.
Try it without writing code
dotnet tool install -g TranscribeCppSharp.Cli
transcribe jfk.wav --model whisper-tiny
The model is fetched from HuggingFace on first use (42 MB for that alias), verified by sha256 and cached; later runs are offline. It runs on the GPU by default.
What it does not do
Stated plainly, so nothing is implied. Each item has its own page with the detail.
- The library is not thread-safe. At most one run per model at a time, across all of that model’s sessions; parallel workers each need their own
Model. This is an upstream 0.x limitation, not one this wrapper imposes (Concurrency). - Streaming cannot attribute speakers.
transcribe_stream_paramshas nodiarizefield in transcribe.cpp v0.2.4, so the streaming API cannot request it; we do not invent one (Diarization). - No CUDA runtime ships in the packages. The bundled binaries are CPU + Vulkan on Windows/Linux and Metal on macOS; an NVIDIA GPU means supplying your own CUDA build (Using CUDA).
- The MIT license does not cover the models. They come from different ecosystems, some non-commercial, and this project neither bundles nor curates them (Model licenses).
- Speaker attribution is not verified in CI. It is covered only when the ~617 MB MOSS asset is fetched with
WITH_DIARIZATION_MODEL=1. Until then, a regression that emptiedSpeakerSegmentswould go unnoticed (diarization). - There is no GPU in CI. The tool is tested end to end on the CPU, but GPU selection itself is covered by policy tests, not by a real run (what the CLI does not do).
Where to go next
| Page | What is in it |
|---|---|
| Getting started | installing, the native packages, load-and-transcribe, batch, streaming, sample code |
| Command-line tool | installing transcribe, a first run, the full --help, and what the tool does not do |
| Compute | GPU by default, --list-devices, backend and device selection, CUDA |
| Models | aliases vs. HuggingFace specs, --list-models, caching, capabilities, model licences |
| Exporting | --out, and the plain, WebVTT and JSON formats |
| Long audio and diarization | windowing, the four diarization entry points, and what is not verified |
| Audio input | which formats are tested, ffmpeg, and the measured memory behaviour |
| Real-time streaming | StreamSession, partial results, its two limits |
| Concurrency | blocking calls, thread safety per type, dispose discipline |
| Architecture | the layers, native library resolution, versioning vs. upstream |
| Error handling | TranscribeException, StatusCode, the failures you will actually meet |
| Development | building, testing, the checks that keep these docs honest, building from source |
| Governance | security, licence, attribution |
Source on GitHub · NuGet · CHANGELOG · MIT