BM25 is only as good as your tokens
Blog post from Neon
Neon’s Lakebase Search has introduced the lakebase_tokenizer extension, which lets managed Postgres users define custom synonym and stop-word dictionaries as SQL tables rather than server-side files. Because BM25 ranking depends on the normalized tokens created before indexing, the extension enables applications to map vocabulary variants such as “k8s,” “kube,” and “Kubernetes,” normalize abbreviations like “PG” and “creds,” and strip accents so searches for “zurich” can match “Zürich.” The configuration works with standard PostgreSQL tsvector columns as well as GIN and lakebase_bm25 indexes, and it is durable and branch-aware, allowing dictionary changes to be tested in database branches. An example using support tickets shows that custom tokenization can increase relevant matches and alter BM25 scores by producing more accurate term-frequency and rarity statistics, while custom stop words can prevent conversational filler terms from causing otherwise useful searches to fail.
| Trend | Post Mentions | Total Month Mentions | Posts | Companies | MoM |
|---|---|---|---|---|---|
| Kubernetes | 22 | No monthly metrics for this publish month. | |||
| Vector Search | 2 | No monthly metrics for this publish month. | |||
Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.