Where developers learn, build, and share. Elastic Dev is your source for hands-on demos, cheat sheets, explainers and more. Check out: @elastic

United States
👋 Welcome to Elastic Dev — your new go-to for hands-on demos, cheat sheets, concept explainers and more. Get the latest on search, data, and GenAI straight from the source. Expect: - Practical demos and how-tos - Posts and technical deep dives from our Labs - Open source updates & more For anyone who loves building, learning and experimenting: let's get started. Hello, world.
14
14
78
50,547
The hard parts of hybrid retrieval, already done. Elasticsearch Vector Database is a new serverless offering where expert-level tuning is the default: - bfloat16 storage: half the disk footprint before quantization even starts - BBQ: up to 32x vector compression, 95% less memory - Auto-calibration re-tunes quantization on every merge as your data drifts - Filtered vector search at up to 8x higher throughput than OpenSearch - Jina AI embeddings and reranking on managed GPU inference, or bring your own models You bring documents and queries. We handle the embeddings, tuning, and infrastructure. Full breakdown, including the semantic_text quickstart and what ships in vectorDB index mode: go.es.io/3Vk5Wcm
2
188
Elastic Dev retweeted
Do you know why BM25 is still useful when we have vector search? BM25 stops rewarding the 20th "cat" in a document. Plain TF-IDF lets term frequency grow without limit. A keyword-stuffed page outranks a good one. Two parameters fix that. k1 caps repetition. The term frequency factor climbs toward an asymptote of k1 + 1, so the 20th hit adds almost nothing over the 5th. b handles length. |D| / avgdl compares a document against the collection average and penalises the long ones.
7
7
61
1,298,138
Elastic Dev retweeted
Before: look up the field name, check the metric type, pick the right aggregation, build the chart. Per metric. Now: type TS metrics-* in Kibana Discover. Discover appends METRICS_INFO to your query behind the scenes. One row per metric: name, type, unit, dimensions, and the data stream it lives in.. Gauges get AVG. Counters get SUM(RATE()). Histograms get PERCENTILE(p). You never choose. Break the charts down by any dimension your data exposes.. Switch data streams and dimensions that don't exist there get cleared automatically. The inventory is your metrics, charted correctly, from one line of ES|QL.
1
3
11
2,237
Elastic Dev retweeted
Write load hotspots can hide inside a single weighted score. Elasticsearch sums 4 shard metrics into 1 score per node. High write load and low shard count cancel out. The node looks balanced, but the hotspot stays put. Elasticsearch Serverless stopped using a single score. Heap, write load, index shard spread: each gets its own decider. A shard only moves for a specific, named reason. One workload cut shard relocations by 50% at the same write throughput. Fleet-wide, data node OOMs dropped to near zero after the heap decider rolled out.. Full technical deep dive, including the production graphs: go.es.io/4ird54v
2
17
2,286
Elastic Dev retweeted
1 question to an AI agent: 24 spans, 10 model calls, 160K input tokens. Elastic 9.5 records all of it out of the box. There's no collector to set up. Every Agent Builder run lands as OTel traces in your own cluster, down to each ES|QL query the agent wrote and which index it ran against. Half the model calls went to a cheaper model. You'd never know from the answer text. It's only in the trace. One thing traces don't hold by default: content. Prompts and responses stay off unless you flip the privacy toggles. And they never tell you who approved the action. That record you write yourself.
3
9
35
1,019,655
Elastic Dev retweeted
Save these copy-paste ES|QL queries for your next Kubernetes incident diagnosis. Blog with all 9 in the reply.
6
2
51
1,082,125
Elastic Dev retweeted
Your ES|QL query dies with Unknown column [field] because an alias got repointed and the new backing index dropped a field. Or you spot a field in your documents that never made it into the mapping. Until now, using it meant a reindex. Hours of it. In Elasticsearch 9.5, SET unmapped_fields="LOAD" reads the missing field straight from _source instead. "NULLIFY" fills the field with nulls. Queries keep working through mapping changes, without touching the data.
6
5
56
2,451,422
Elastic Dev retweeted
470 of the 500 slowest requests share 1 log pattern. You'd never find it reading traces one at a time. At 10 traces the check is tedious. At 500 it doesn't happen. ES|QL subqueries are now on Serverless and in tech preview in 9.5 WHERE trace_id IN (subquery) The trace query runs inside Elasticsearch. The outer query pulls every log line those requests wrote, across every service, and groups them into log patterns. In this investigation it surfaced a lock wait timeout 2 hops downstream from the alert, in a service nobody was watching. 1 slow trace gives you a theory. 470 of 500 tells you which team to page.
1
3
9
2,289
Elastic Dev retweeted
Elasticsearch 9.5 cuts timestamp storage by 92%. 1.03 GB down to 79 MB on a high-cardinality benchmark of 2.26 billion data points. The old time series codec applied one fixed encoding to every numeric field. Fine for timestamps and counters. Useless for floats: a 0.01 change in a CPU gauge is a jump of trillions in integer space. ES95 selects the encoding per field instead, based on what the mapping already declares: - SplitDelta: splits blocks at series boundaries, so one jump stops taxing the whole block - ALP: recovers the decimal structure in floats, compresses them as integers Per-field results: gauges down 19% to 74%, counters 20% to 30%, total doc values down 33.6%. Dimension fields like host.ip aren't covered yet; that work is ongoing. Applies to newly written data on upgrade. Existing segments keep their encoding, no migration required.
10
15
108
2,072,322
Elastic Dev retweeted
"It feels slow" is the worst ticket a self-hosted LLM generates. vLLM starts, serves, and reports success whether it's configured brilliantly or wastefully. Nothing tells you which. But the answer is already sitting on /metrics: latency split by phase, cache hit rates, batch occupancy, completion outcomes. Almost nobody scrapes it. On a single A10G, 6 metric checks turned "it feels slow" into a verdict: TTFT: 51 ms against a 300 ms p95 target Queue time: 0.01 ms Decode: 97% of total latency Prefix cache: 32% of prefill work skipped entirely The server was over-provisioned, not under. The slowness lived somewhere in front of it: the app, the gateway, or the prompt. Tuning an LLM server is reading telemetry and reasoning about saturation. SREs have been doing that for 20 years. Full walkthrough with the ES|QL queries, the DCGM setup, and all 6 metric checks on an A10G: go.es.io/3U9DTfm
4
3
11
2,590
Elastic Dev retweeted
AI Where It Counts: Relevance Please - Episode 6 nitter.net/i/broadcasts/1nJOLQbAq…
1
2
8
2,214
Elastic Dev retweeted
Code is the only documentation that's never outdated. Our field teams get asked things like: does App A v1.2.3 work with App B v9.8.7 on Kubernetes? Docs rarely cover that. The code always does. So one of our field engineers built Sourcerer. It searches a billion lines of our code across every repo and version we support. Every answer cites the exact file and line. Built on Elasticsearch and Agent Builder. Open source, Apache 2.0. Full benchmarks and solution design: go.es.io/4xMJpmq
2
6
14
3,085
Elastic Dev retweeted
🧵 jina-embeddings-v5-omni represents the next evolution of embedding models. Literally. The vision pipeline uses the Qwen3.5 vision encoder, which itself is based on SigLIP2. Audio comes through the Qwen2.5-Omni audio encoder, built atop Whisper-large-3. Text runs through jina-embeddings-v5-text. So how do we get all of these different modalities, all running through different encoders, into one embedding space? The architecture that makes it work is called GELATO.
1
4
12
2,407
Elastic Dev retweeted
Your agent burns tokens just figuring out where to look. It inspects mappings, samples docs, probes indices. That's 12 tool calls and 167K tokens to get an answer. Knowledge Indicators fix this. A Kibana Workflow profiles each index once: purpose, key fields, routing heuristics. Those profiles live in an AI Index, and agents query them with ES|QL. Same question: 8 calls, 92K tokens, correct answer. The context was already there.
2
2
15
2,605
Elastic Dev retweeted
When you create a recipe, you try a few tweaks and combinations before landing on the final version. Creating an embedding model is no different. This process of experimentation is known as ablation. For both audio and vision pipelines in jina-embeddings-v5-omni, multiple ablation studies were performed before we landed on the final architecture: freeze all the encoders, train only the smaller projectors.
7
6
48
2,092,840
Elastic Dev retweeted
Relevance Please: Ep 5 - Build Agents That Ask Before They Act nitter.net/i/broadcasts/1rGmqpppE…
2
1
15
2,272
Elastic Dev retweeted
How can a text embedding model map audio and image vectors? After being transformed via a projector (translator), a foreign modality vector must announce itself before entering an embedding model for an unrelated modality. Think of HTML. <p> These tags let the document know this element is a paragraph. </p> Vectors work the same way.
1
4
17
2,440