How matching decides that two names mean one thing

Four sieves in rising cost, two thresholds we tuned by hand, and why Postgres and PostgreSQL merge while MySQL and PostgreSQL do not — although a vector can barely tell the two pairs apart.

4 min


Everything in this system rests on one decision: when do two names mean the same thing? Get it wrong in one direction and two people who know the same thing sit on separate nodes, matching neither each other nor the same job. Get it wrong in the other and a candidate is credited with knowledge they never claimed.

Here is how it is decided, with the numbers.

Four sieves, in rising cost

A name arrives — from a conversation, a CV, a job listing. It goes through four steps, cheapest first, and stops at the first that answers.

1 · The exact key. The name is normalized — case folded, punctuation stripped, whitespace collapsed — and looked up. This answers most of the time, and it costs a single indexed read.

2 · A known alias. Names we have already learned belong to a concept. k8s → Kubernetes. The important detail: an alias is stored both as written and as its normalized key, because the lookup reads the key. Writing only the display form makes the alias decorative — it looks present in the database and is never found. We made exactly that mistake once.

3 · Meaning. An embedding comparison against existing concepts, weighed together with how close the names are as strings. This is the expensive step and the interesting one.

4 · A new concept. Nothing matched, so the vocabulary grows by one.

The two thresholds

The merge decision is not a single number, and this is the part worth copying if you are building something similar.

  • 0.80 by vector alone. Two names whose embeddings are this close are the same thing, whatever they look like.
  • 0.45 by vector together with 0.55 lexically. Moderately close in meaning and close as strings.

Why both? Consider two pairs:

PairVector similarityLexical similarityMerge?
Postgres ~ PostgreSQLhighvery highyes
MySQL ~ PostgreSQLhighlowno

By vector alone these two pairs are nearly indistinguishable — both are "relational database" shaped, and an embedding model is measuring topic, not identity. The names are what separate them. A system using only embeddings merges MySQL into PostgreSQL and then tells a company that a candidate knows something they do not.

Why we would rather miss a match than invent one

The thresholds err toward missing. A missed match costs a candidate one listing they might have been right for, and they will see the unmet requirement and can say so. An invented match costs a company an interview with somebody who cannot do the job, and costs that candidate the interview too — and nobody finds out until it is too late to be useful.

That asymmetry is specific to hiring. In a recommendation engine a false positive is a wasted scroll. Here it is a wasted day, on both sides, built on something nobody said.

Changing the thresholds without losing anything

Interpretation — which concept a claim points to — is the third of three layers and the only disposable one. The raw source never changes; the claim with its sentence and its time is only ever added to. So when a threshold moves, the interpretation version is raised and layer three is rebuilt from layer two.

Everybody's graph improves. Nobody's evidence is touched. And a bad threshold is a bad afternoon rather than a lost database.

There is one rule that makes this safe, and it is easy to break: any field that is going to survive a rebuild has to exist on the claim, not only on the interpretation. A field that lives only in layer three disappears the next time layer three is rebuilt — silently, because a rebuild is supposed to replace it.

The trap that cost us a ranking

A footnote for anyone moving between vector stores. pgvector's <=> operator returns cosine distance — that is 1 - cosine — so similarity is 1 - (a <=> b). The graph database we used before returned (1 + cosine) / 2, and the old queries wrote 2 * score - 1 to correct for it.

Both forms are right in their own place. Mixing them inverts the ranking with no error at all: the query runs, the numbers look like numbers, and the worst match sorts first. It broke matching silently once before anyone noticed.


Related: how a match is computed · how it works, end to end

Read next

Check it yourself.

Everything described in this article is visible in the product: the sources, the claims with their sentences, and the match with its working.