Query Result Scoring

Every path: / namespace: / nodeType: / source: query in the mesh flows through MeshQuery. It fans out to every registered IMeshQueryProvider, collects their results, and emits a single sorted QueryResultChange<T> to the caller. This page explains how that merge orders results — the contract each provider must follow, and the sorting rules the aggregator applies. MeshQuery fan-out PostgreSQL prefix 100 · sub 50 · prox 40 StaticNodeQuery FuzzyScorer / unscored Custom Provider Scores[ ] or null ClipMergedInitial 1. EffectiveOrderBy 2. Score desc 3. Path tiebreak Sorted Results Skip / Limit select: projection QueryResultChange Each provider scores independently; ClipMergedInitial owns the cross-provider sort.

Query fan-out, per-provider scoring, and aggregated sort pipeline.

Result Shape

QueryResultChange<T> carries the following fields:

Field Purpose
Items The result items (typically MeshNode).
Scores Optional parallel array — one double per item. Higher = stronger match.
Query The parsed query, giving the aggregator access to OrderBy.
Version, Timestamp Bookkeeping for change feeds.

When Scores is null, the aggregator pairs every item in that batch with 0.0. Because OrderByDescending is stable, an all-unscored result keeps its insertion order — but against a provider that did score, an unscored batch competes as score 0 and lands below any positive hit. When non-null, its length must equal Items.Count. Each provider independently decides whether to score its results.

Sort Dimensions

MeshQuery.ClipMergedInitial is the authoritative sort pass. It runs after every provider's Initial emission has arrived and before Skip / Limit trim the window.

Dimensions are applied in this order:

  1. ParsedQuery.EffectiveOrderBy (when it resolves). The author's sort: when one was written — intent always wins — and otherwise, for a query with no free-text term and the default source:, ParsedQuery.DefaultFilterOrdering: lastModified descending, newest first. A query like ... sort:LastModified-desc sorts by MeshNode.LastModified descending via QueryEvaluator.OrderResults; a query like nodeType:InstanceAction limit:25 sorts the same way without saying so. ParsedQuery.OrderBy keeps meaning "what the author asked", so a surface can print the applied ordering and say whether it was authored or defaulted.
  2. Score descending. When no ordering resolves — a free-text term, or a change feed (source:activity / source:accessed, ranked by their provider) — score is the sole sort key. Highest score lands at index 0.
  3. Path ascending as the final tiebreaker (PathTiebreak). This is what makes the order TOTAL, which is what Skip/Limit need to be paging rather than sampling — insertion order, which this used to name, is the order two providers' emissions happened to merge in, and even one provider's scope walk emits in read-completion order.

After sorting, Skip and Limit clip the window. The select: projection runs last — projected dicts and anonymous types are emitted only at this boundary.

Why a filter-only query defaults to newest first

A capped result is a partial answer, and which part survives the cap is decided by the order it was clipped over. A filtered query has no relevance signal, so before the default each clip site took its window over whatever its backend enumerated: path-alphabetical on the in-memory walk, heap order on Postgres (no ORDER BY before LIMIT). The failure that motivated the rule (MeshWeaver #4950): two truncated: true pages whose newest row was six days old, over a set whose actual newest rows explained the outage — the rows were simply not in the first 25, and nothing in the response distinguished "these are the newest" from "these are the first the heap gave up". Ordered newest first, a truncated page is still partial but it is the RIGHT part; unordered, it is a misleading answer that looks complete. This is the same rule as the coverage denominator: a zero is read against coverage.partitions (Search Coverage and Refusal), and a cap is read against the order.

The default is resolved ONCE, on the contract, and read at EVERY clip site. StorageAdapterMeshQueryProvider clips to Skip + Limit before the merge ever sees a row, and so does the Postgres generator's LIMIT; a default applied only in ClipMergedInitial would re-sort a subset those sites had already chosen in the old order. So the in-memory provider and the merge both read EffectiveOrderBy, and the Postgres ORDER BY branch (MeshWeaver.Plugins, PostgreSqlSqlGenerator) reads the same property once it moves to a platform pin that carries it — until then a Postgres portal's filter-only page is still heap-clipped and merely re-sorted newest first by the merge, which is the difference between "an ordered partial answer" and "the right partial answer". FilterOnlyQueryClipsNewestFirstTest pins both core halves, and its controls pin what the default must not touch: an authored sort:, a free-text term.

Per-Provider Scoring Conventions

Cross-provider comparability is the key invariant. Score scales must be comparable across providers for the same query. A PostgreSqlMeshQuery name-prefix hit (score 100) should beat a StaticNodeQueryProvider plain-listing hit (which emits no score, so it competes as 0) when the same query reaches both providers.

StaticNodeQueryProvider

Source: src/MeshWeaver.Hosting/Persistence/Query/StaticNodeQueryProvider.cs

Query shape Score
Text search (textSearch:foo or free-text tokens in the query) FuzzyScorer.Score(...) against the node's Name (falling back to Path) — fzf-style: ~16 points per matched character plus boundary / consecutive / camelCase / after-separator bonuses. Not normalised — the magnitude scales with query length, so a typical 5–10 character term lands in the low hundreds. Non-MeshNode items score 0.
Filter / namespace / nodeType only Scores = null — no relevance signal to surface ("give me all Threads in this namespace" is unordered with respect to score). The aggregator then treats every item as 0.

StorageAdapterMeshQueryProvider

Source: src/MeshWeaver.Hosting/Persistence/Query/StorageAdapterMeshQueryProvider.cs

This provider does not score. Every Initial it emits leaves Scores unset (null), whatever adapter is wrapped — it uses a fuzzy score internally for its own autocomplete/suggestion ordering, but never publishes one on the QueryResultChange. Its items therefore enter the merge at 0 and keep their insertion order among themselves.

PostgreSqlMeshQuery and PostgreSqlPartitionedMeshQuery

Sources: MeshWeaver.Plugins/src/MeshWeaver.Hosting.PostgreSql/PostgreSqlMeshQuery.cs, MeshWeaver.Plugins/src/MeshWeaver.Hosting.PostgreSql/PostgreSqlPartitionedMeshQuery.cs, MeshWeaver.Plugins/src/MeshWeaver.Hosting.PostgreSql/PostgreSqlSqlGenerator.cs

There are two distinct rankings in the PostgreSQL layer; do not conflate them.

(1) The published Scores[] are computed in C#, by PostgreSqlMeshQuery.ComputeRowScores, over the rows the initial emission carries:

Component Score
Name prefix match 100 - (name.Length - termLength) — shorter prefix-matched names rank higher
Name substring match 50
Path substring match 30
Path proximity boost PathProximity.ComputeBoost(contextPath, resultPath) — 40 / (1 + segmentDistance), so max 40, decaying with namespace segment distance (src/MeshWeaver.Mesh.Contract/Query/PathProximity.cs)

The three text buckets are mutually exclusive (first match wins — prefix, else substring, else path); the proximity boost is then added. ComputeRowScores returns null — i.e. no scoring at all — when the item is not a MeshNode, or when the query has neither a text term nor a context path, so a purely structured query is deliberately left unranked rather than amplifying a constant 0.

(2) A separate SQL-side relevance ladder decides which rows survive LIMIT, on its own scale (exact name 1000, name-prefix 600, id-prefix 500, name-substring 300, id-substring 200, description-substring 100) in PostgreSqlSqlGenerator. It exists so the database keeps the most relevant rows when it clips, before the C# merge ever sees them. It is not what lands in Scores[].

Vector search is a third, separate ordering: GenerateVectorSearchQuery orders by cosine distance (n.embedding <=> @queryVector, with a lexical tier in front when a term is present) and projects _distance. Cosine similarity is not folded into Scores[] — rows returned by the vector path are re-scored by ComputeRowScores like any other. See Vector Search.

Adding a New Scored Provider

To hook into the aggregator's ranking:

  1. Compute one numeric score per result item inside your provider.
  2. When building the Initial QueryResultChange<T>, set Scores = items.Select(ComputeScore).ToList().
  3. Choose a scale that won't be drowned out by the PostgreSQL bonuses (100 / 50 / 30) when the same query reaches both. If you can't reasonably rank, set Scores = null.

Why the Aggregator Owns the Sort

A single provider can rank within its own result set, but cross-provider tie-breaking requires a single decision point. A PostgreSQL hit with name-prefix score 100 must beat a static-catalog hit with score 0, even though both Initial emissions arrive independently. Placing the sort in ClipMergedInitial guarantees that every downstream consumer of Query<T> / QueryAsync sees the same deterministic ranking regardless of which providers contributed.

Legacy: The "Writable First, Static Last" Ordering

Before the current scoring contract, MeshQuery.MergeProviderObservables ordered provider buckets as writable-persistence first, static catalog last to prevent static entries from crowding out user content under a limit: clause. That heuristic was a stand-in for proper scoring. With per-provider Scores it is gone — the merge is now a flat concat with no priority shuffle: PostgreSQL sets a high score for relevant rows, the static catalog leaves Scores null for filter-only matches (so its items enter at 0), and Limit clips exactly the right tail.

See Also