Observability

Three questions, three tools. Why did this rank here? — for one search, per document: SearchTrace below, and Explain in Ranking. How long, how many candidates, which stage is slow? — across all searches in production: RetrievalTelemetry. Does the whole fleet agree, or is one retriever insisting? — for a fused page: RetrievalAgreementAnalyzer. Plus a score-to-confidence mapping for gating a result set you are not sure about.

All of it is opt-in and dependency-free. A null telemetry, or a search without a trace, allocates and records nothing.

Ranking trace (SearchTrace)

Explain breaks down one scorer’s arithmetic. SearchTrace goes further: it records every stage a document passed through, so a final rank can be walked back stage by stage — which engine scored it, what each merger contributed, what the reranker changed, what the boost did, and which route ran.

using LexiSharp.Core;

var trace = new SearchTrace();
var hits = engine.Search("refresh token", new SearchOptions(Limit: 5, Trace: trace));

foreach (var step in trace.Steps)
    Console.WriteLine($"{step.Stage,-6} {step.DocumentId,-8} {step.Before,8:0.000} -> {step.After,8:0.000}  {step.Detail}");

// route   d1              0.800 ->            dense (fallback: below threshold or no opinion)
// score   d1      4.819997 ->      4.819997  BM25
// merge   d1      0.016393 ->      0.016393  lexical=4.819997 dense=0.812003
// rerank  d1      4.819997 ->      0.912000
// boost   d1      0.912000 ->      1.824000  x2 +0

SearchOptions.Trace defaults to null, and a null trace is inert: no engine records, allocates or formats anything. Pass one and each stage appends to it as it runs, so a decorated pipeline accumulates its whole chain.

Cost — what is measured and what is not. Measured with BenchmarkDotNet’s MemoryDiagnoser and stable across repeated runs: a saturated trace allocates 0 additional bytes per search (1.34 KB either way), and a trace created per search costs ~0.47 KB (the object plus its backing array).

A time figure is deliberately not given, and the reason is itself a measurement. On the machine used for the run the baseline benchmark varied between 1.52 ms and 3.57 ms (2.3×) across identical runs, BenchmarkDotNet reported bimodal distributions on the pre-existing search benchmarks too, and the traced variant came out faster than the baseline in one run — which is impossible. A delta under that spread is not measurable there, so quoting one would be fiction.

Full details, machine configuration and the reproduce command: Benchmarks.

Bounds. A trace stops recording past Capacity (default 256) and counts the overflow in Dropped; check IsTruncated rather than assuming Steps is the whole story. Stage recording is bounded by the page or the candidate shortlist, never by the corpus, so a trace cannot grow with the index — this is what the ScoreStageIsBoundedByThePageNotTheCorpus test pins on a 20 000-doc index.

Not thread-safe. A trace is a mutable collector: give each concurrent search its own, the same way each gets its own SearchOptions. The telemetry below is the opposite case — immutable and stateless per query, so one instance is shared.

Observability (RetrievalTelemetry)

SearchTrace explains one search. RetrievalTelemetry answers the operational question — over time, in production, across all searches: how long they take, how many candidates they touch, how big the index is, and which stage is responsible when a query is slow.

It is opt-in and dependency-free. Attach a telemetry to an engine and it reports to a RetrievalLog sink, an IRetrievalMetrics collector, or both:

using LexiSharp.Core;

var metrics = new InMemoryRetrievalMetrics();

var engine = new RankedTextSearchEngine(
    index, new Bm25Scorer(),
    telemetry: new RetrievalTelemetry(metrics: metrics));

engine.Search("refresh token");

foreach (var e in metrics.Snapshot().Engines)
    Console.WriteLine($"{e.Engine}: {e.SearchCount} searches, {e.MinElapsedMs:0.###}–{e.MaxElapsedMs:0.###} ms");

With LexiSharp.AspNetCore, an ILogger sink is one call away:

using LexiSharp.AspNetCore;

var telemetry = RetrievalLogger.Telemetry(
    loggerFactory.CreateLogger("search"), RetrievalLogLevel.Information);

Events carry the engine name, a stable event id (search / stage / index), a self-contained message and structured ElapsedMs / Count fields, so a log aggregator can build dashboards without parsing prose. Per-stage events are Debug, so the default Information threshold keeps one line per query and drops the per-stage detail.

Instruments by pipeline stage. HybridTextSearchEngine reports one source stage per delegate (labelled by your sourceNames, so a slow lane is named the way your application named it) and one merge stage. RerankedTextSearchEngine reports retrieve and rerank:<name> — the cross-encoder call is usually where the latency lives, and it warns above 100 ms. BoostedTextSearchEngine reports retrieve and boost. A query against an empty index, or one whose terms match nothing in the vocabulary, is a Warning rather than a silent zero-result page.

Cost when unused. RetrievalTelemetry.None reports IsEnabled == false, and every instrumented engine reads that once per search before touching a timer, so an un-instrumented engine takes no timestamp, allocates nothing and calls nothing. A RetrievalTelemetry is immutable and holds no per-query state, so concurrent searches share one safely as long as the sinks are. Note this is a statement about the code, not a measurement: no timing figure is quoted here for the same reason the trace section above quotes none.

What is not covered. IRetrievalMetrics is the only number surface: there is no built-in OpenTelemetry or Prometheus exporter, and no percentile histogram — InMemoryRetrievalMetrics keeps a count, a total, a min and a max per engine and stage, which is enough to assert on in tests and to back a diagnostics endpoint, and is deliberately not a substitute for a real metrics pipeline. Implement IRetrievalMetrics over your own backend for that. SearchTrace.StageCounts() reports how many steps each stage recorded; per-stage timings come from the metrics collector, not the trace.

A LexiSharpIndexHealthCheck reports readiness for the same reason: an index holding nothing answers every query with zero results, which is a silent failure rather than an obvious one.

services.AddLexiSharpRetrievalMetrics();
services.AddLexiSharpSearchHealthCheck(
    _ => index,
    o => o.Tags = ["ready"],
           o.Probe = ct => connection.OpenAsync(ct));  // optional: backing-store reachability

Source agreement (RetrievalAgreementAnalyzer)

A fused ranking hides why a document is on the page. RetrievalAgreementAnalyzer reads the per-source scores in DetailedSearchResult.Contributions and classifies each document by how much its sources agree:

using LexiSharp.Core;

var page = hybrid.SearchWithDetails("refresh token", new SearchOptions(Limit: 20));
var reports = RetrievalAgreementAnalyzer.Analyze(page);

foreach (var report in reports)
    Console.WriteLine($"{report.DocumentId}  {report.Agreement}  strong: {string.Join(",", report.StrongSources)}");

var counts = RetrievalAgreementAnalyzer.Summarize(reports);
Category What it means
Unanimous several sources returned it, all found it convincing
Disputed several returned it, they disagree — some convinced, some not
Lukewarm several returned it, none convinced: it is on the page through a merger, not anyone’s conviction
SingleSource exactly one source returned it — the case worth surfacing, since a reader of the final score cannot tell
None no source returned it; should not occur on a result page

The categories are structural, not named, because source labels are yours: a setup calling them lexical/dense and one calling them bm25/cosine describe the same shape, so a hard-coded LexicalOnly would be wrong for half of them. StrongSources gives the names back.

Normalization is not optional. Raw contributions are not comparable across sources — LexiSharp’s dense lane returns a cosine in [0, 1] while its BM25 lane returns values around 1 to 10. Dividing each contribution by the best score that source gave on the same page makes the values scale-free without assuming anything about the scoring function. A source that returned nothing is reported as AbsentSources rather than weak, because being outside a source’s depth is a different fact from being ranked low by it.

The threshold is a heuristic. strongThreshold (default 0.5) is the share of a source’s best score at which a document counts as strongly supported. It is a round number chosen for readability, not a value validated against relevance data — tune it against your own corpus, and pass sourceNames so a source that silently contributed nothing still shows up as absent.

The demo is the quickest way to see the distinction matter. Running its six sample queries over its 26-document corpus, the fused page holds 9 unanimous, 8 disputed, 7 singleSource and 2 lukewarm — the obvious matches are unanimous, the arguable ones disputed, the lukewarm ones are on the page through the merger with nobody enthusiastic, and some ride on a single lane. None does not appear, and cannot: it classifies a document no source returned, so by definition it never shows up on a result page.

Calibrated confidence (ScoreConfidence)

BM25 output is not a probability: a score of 4.8 means nothing on its own, and the scale differs per scorer. ScoreConfidence maps a result set’s raw scores to one confidence in [0,1] per result, so a result set can be gated by an application-level minConfidence threshold that raw lexical scores cannot express:

using LexiSharp.Ranking;

// Per result, in ranking order. WinnerMargin is the default.
IReadOnlyList<double> confidences = ScoreConfidence.Compute(results);

// The one to compare against a gate. 0 for an empty result set.
double top = ScoreConfidence.TopConfidence(results, ScoreConfidenceMethod.ZScore);
Method What it measures Blind to
WinnerMargin (default) the gap to the next result, 1 - score_next / score_current; the last result is 0 the absolute quality of the match — a flat ranking of high scores still looks unsure
ZScore each score’s position in the set’s own distribution, squashed through the logistic function; an all-tied set collapses to a neutral 0.5 whatever its scale the scale itself, so two corpora are not comparable

A score of exactly 0 means not a match by engine convention, so such results — and any non-finite score — are never confident. Read the name carefully: this is a relative confidence in a result set, not a calibrated probability of relevance. The threshold is yours to choose, and nothing here was fitted against relevance judgements.


This site uses Just the Docs, a documentation theme for Jekyll.