Open-source AI orchestration framework by @deepset_ai. Build context-engineered agents & RAG systems in Python. Discord for support → discord.gg/AH8a5Tm8vb

Based in Germany
Haystack 3.2 is here 🚀 This release keeps agents lean and in control as conversations grow. Three highlights: 🧵 SummarizationCompactor for smarter context compaction Pair SummarizationCompactor with CompactionHook (shipped in v3.1) to progressively summarize long-running Agent conversations. 💰 TokenBudgetHook caps what a run can spend A ready-made hook that ends the run at "token_budget_exceeded" once you hit your token ceiling through the stop_run hook point. 🛠️ Faster Pipeline building New add_components() and connect_many() let you add and wire up multiple components in one call. 💙 70 contributors made this release happen; huge thanks to everyone! 🔗 Full release notes: haystack.deepset.ai/release-…
4
1
10
247
Full-text search without semantic understanding leaves gaps. Dense vector search without keyword matching misses precision. @ApacheSolr is now available as a Document Store and retriever in Haystack. Use SolrDocumentStore with SolrHybridRetriever to combine BM25 full-text search and dense vectors in one pipeline - getting both lexical and semantic matching without extra infrastructure. Solr has been a trusted open-source search engine for over a decade, and its hybrid retrieval approach fits naturally into Haystack pipelines whether you're prototyping locally or scaling to production. 🐍 pip install solr-haystack 🔗 haystack.deepset.ai/integrat…
7
7
217
📣 We're hosting our next unconference in Berlin with @prior_labs, exploring open-weight models and open-source agents! Come join us if you'd like to discuss: → Structured data in agentic workflows → When agents should hand off to TabPFN for prediction → Benchmarking tabular models vs. LLMs → Why orchestration matters for production agents → What owning your stack end-to-end actually requires → Open-weight vs. closed models in practice Or anything else that's been challenging you lately! 📍 @deepset_ai HQ, Berlin 🗓️ September 24th, 6PM-9PM Register: luma.com/haystack-priorlabs
2
1
8
248
If you are a @mariadb user, you can now store your embeddings right alongside your documents and metadata in the same system - no need for a separate vector database. MariaDBDocumentStore and MariaDBEmbeddingRetriever let you build semantic search pipelines that keep all your data unified and easy to manage. Embed your content, persist vectors in the same MariaDB instance where your source data lives, and retrieve similar results in seconds. Everything stays in one place. 🐍 pip install mariadb-haystack 🔗 haystack.deepset.ai/integrat…
6
231
Haystack 3.1 is here 🚀 This release is all about smarter context management with token accounting, context compaction & precision. Three highlights: 🪝 Context compaction for agents New CompactionHook automatically shortens long-running conversations before they blow through the context window. 🪙 New token counters for estimating usage Three new counters, from a zero-dependency approximation to an exact, OpenAI API-backed count, so you know when to compact before you hit the wall. 🛠️ AgentTool for multi-agent systems Wrap any Agent as a Tool so another Agent can delegate to it. Only the final reply crosses back, the coordinator's context stays clean. 💙 42 contributors made this release happen; huge thanks to everyone! 👉 Full release notes: haystack.deepset.ai/release-…
1
1
13
350
HetznerChatGenerator brings European sovereign cloud LLMs to Haystack. Process text and images through @Hetzner_Online infrastructure while keeping your AI workloads within EU borders. 🐍 pip install hetzner-haystack 🔗 haystack.deepset.ai/integrat…
3
228
We are thrilled to be among the ten finalists for The Spark – The Technology Award. For eleven years, @handelsblatt and @McKinsey have been recognizing Germany’s most innovative technology startups. This year’s theme is "Scaling Deep Tech", focusing on technologies that are already making a real-world impact today and shaping the future of industry and the economy. Now, we need your support. Vote by September 7th to help us reach the Top 3! Your vote counts directly toward the jury's decision. Vote for us here: cmk.handelsblatt.com/cms/art…
2
164
Finally, @Linkup_platform joins Haystack's growing search component library. Fetch real-time web results, get ranked documents with source links, and integrate them into your pipeline in a few lines. Same philosophy as other Haystack components - simple, direct, pluggable. haystack.deepset.ai/integrat…
3
180
📷 DDGS (Dux Distributed Global Search) is now available in Haystack. Query Google, Bing, Brave, Yahoo, Yandex and more from your pipelines and agents. Multi-source web search without the infrastructure overhead. pip install ddgs-haystack haystack.deepset.ai/integrat…
1
4
211
Please welcome Mirage by @struktoai to the Haystack integrations family. Create isolated virtual filesystems with multiple mounts, such as in-memory storage backed by cloud resources like S3, and give your agents the ability to explore and analyze files using bash commands without exposing your infrastructure. Agents can now run grep, cat, head, wc, and other shell utilities against mounted data sources, treating them as a unified filesystem. Control which commands are allowed and keep sensitive resources read-only while giving agents just the access they need. 🐍 pip install mirage-haystack 🔗 haystack.deepset.ai/integrat…
1
1
5
239
Haystack 3.0 shipped this month. We spent launch week showing you what it can do. Now here's what it took to build it 👇 Since @Haystack_AI 2.0: 2,400+ pull requests merged, 195 contributors, ~170 of them making their first-ever contribution along the way. 17,258 CI runs since the v3 branch was cut, roughly one every five minutes, around the clock, for eight weeks. Most-edited file of the whole cycle: agent.py. Of course it was. Migrating to 3.0? The full guide covers all breaking changes: docs.haystack.deepset.ai/doc… Or skip straight to haystack-v2-to-v3, an agent skill that migrates your v2 pipelines and agents for you. 💙 Huge thanks to our open-source team at @deepset_ai, and to our community contributors who shipped fixes and features for v3 alongside us! Haystack 3.0 is yours as much as it's ours. Got questions about migrating, or curious why we made a specific call? Join us at Haystack 3.0 Office Hours next week, August 4, with @LukawskiKacper and @bilgeycl 👉🏼 Register: luma.com/haystack-3
5
10
492
Day 5 of Haystack 3.0 Launch Week: Human-in-the-Loop from Terminal to Production For today's drop, we implemented a production agent that pauses before anything risky and waits for a human to say go. The approval shows up right where the conversation is happening, making the process intuitive for the users. We built the whole thing as a real, deployable service (Hayhooks + @Redisinc + @OpenWebUI), so it's not just another notebook trick or a simple terminal demo Set up with one command and run all services with @Docker 🐳 🎥 Watch it in action ↓
1
4
9
255
Deep Research Agent
1
4
83
Advanced RAG Agent
1
5
57
Agent Pack ships new with @Haystack_AI 3.0: a collection of complex, pre-configured Haystack agents you can run as they are, customize, or copy as a blueprint for your own architecture. 📬 On Day 3/5 of Launch Week, we're highlighting these pre-built agents: - Advanced RAG Agent: a metadata-aware RAG agent that constructs its own filters to narrow retrieval - Deep Research Agent: a multi-agent system that writes a cited Markdown report, using @tavilyai under the hood for search Two agents today, more landing in the pack over time. Let's see how they work under the hood 🧵
3
3
17
600
Your agent doesn't check with you before it makes its next move. It just runs. That's the agent loop: LLM call, tool call, LLM call, tool call, however many times it takes to finish. Most frameworks let you watch it happen. Very few let you step into it. Day 2 of Haystack 3.0 launch week is about controlling the agent loop for token cost, not just observing it. Every `Agent.run()` already returns `step_count`, `token_usage`, and `tool_call_counts`. Agent hooks let you act on that metadata from inside the loop itself: → Soft limit warns, hard limit blocks → `before_llm` stops an over-budget run before its next LLM call → `after_run` gives you a final report every time, clean, warned, or stopped mid-loop 🎥 Watch it in action ↓
4
4
13
371
Haystack 3.0 is here 🚀 Agents move to the center of the framework: 🤖 Pre-built agents: deep research + advanced RAG, ready out of the box 🪝 Hooks to control the agent loop 🧰 First-class skills with progressive disclosure 📉 A leaner core: we deleted 3 lines for every 1 we added And this is just Day 1. New drops every day this week at 3 PM CEST. Full announcement 👇 haystack.deepset.ai/blog/hay…
10
10
28
1,980
3 days until Haystack 3.0 🚀 Monday, 3 PM CEST | 9AM EST. Then a new drop every day, all week. The countdown is running ⏳
4
1
9
518
Integration alert! @OrcaRouter is now available as a ChatGenerator in Haystack. Use OrcaRouterChatGenerator to swap between standard LLMs, and multimodal models through a unified interface - no pipeline rewiring needed. It acts as a flexible proxy layer that handles tool routing, conversation, and multimodal inputs seamlessly, letting you experiment with different models or providers without changing your agent code. 🐍 pip install orcarouter-haystack 🔗 haystack.deepset.ai/integrat…
1
2
5
665
Your @googledrive is now queryable in Haystack. Use GoogleDriveRetriever to search files stored in Google Drive, then GoogleDriveFetcher to retrieve their content. Built on the OAuth integration, it handles secure token management automatically so your pipelines stay authenticated as users interact with your application. This integration is ideal for RAG applications that need to pull knowledge from shared team documents, spreadsheets, or presentations stored in Google Drive without duplicating data into a separate vector database. 🐍 pip install google-drive-haystack 🔗 haystack.deepset.ai/integrat…
1
6
259
Big news: Haystack 3.0 ships Monday, July 20 🚀 It's one of the biggest releases in Haystack's history, and one announcement wouldn't do it justice. So we're running our first-ever Launch Week: July 20 to 24, five days, five drops, something new every day at 3 PM CEST 🎉 Day 1 kicks things off with the release itself and the full story of what changed and why. After that, each day goes deeper into a different part of 3.0. We won't spoil the lineup, but if you've been following Haystack on GitHub recently, you might see some of it coming 👀 Check the countdown and sign up for updates 👇 haystack.deepset.ai/launch-w… See you Monday 💙
4
7
928
Search and fetch @SharePoint documents directly in Haystack with the microsoft-sharepoint-haystack integration. MSSharePointRetriever queries across your SharePoint sites while MSSharePointFetcher downloads full document content, both powered by the OAuth integration for automatic token management. Perfect for RAG pipelines that need to tap into enterprise document libraries stored in @Microsoft365 without friction. 🐍 pip install microsoft-sharepoint-haystack 🔗 haystack.deepset.ai/integrat…
6
173
Secure OAuth flows are now built into Haystack. The OAuthTokenResolver handles token exchange automatically - connect to Microsoft Graph, Google APIs, and any OAuth provider without manual credential management. OAuth is the standard. It powers authentication across the web. Now your Haystack pipelines can leverage it natively. 🐍 pip install oauth-haystack 🔗 haystack.deepset.ai/integrat…
4
206
Haystack 2.31.0 is here 🚀 This is the final 2.x release before Haystack 3.0. It focuses on preparing the ecosystem for the next major version while delivering quality-of-life improvements today. Many interesting updates this time, but only two highlights.
1
3
4
292
🧠 @cognee_ is now available as a memory store integration in Haystack. Seed permanent knowledge into your graph and then retrieve relevant memories in agentic pipelines. With Cognee and Haystack, you can: - Write domain knowledge once to the permanent knowledge graph - Build agents that combine conversation history with long-term facts - Create session-scoped stores with just a session_id string This makes it simple to build stateful agents that remember facts across sessions and grow smarter as they interact, combining real-time conversation with persistent domain knowledge in a single unified system. 🐍 pip install cognee-haystack 🔗 haystack.deepset.ai/integrat…
2
7
311
Another integration alert! Use FunASRTranscriber to convert audio files to Documents with fully self-hosted speech recognition supporting 50+ languages. No API keys required. FunASR handles transcription, speaker diarization, and timestamp extraction. Pair it with Haystack pipelines to build audio-to-document workflows that stay entirely on your infrastructure. Models are downloaded on first use and cached locally. @ModelScope2022 🐍 pip install funasr-haystack 🔗 haystack.deepset.ai/integrat…
1
5
211
ArangoDB is available as a Document Store in Haystack. Use ArangoDocumentStore to connect directly to your ArangoDB instance, and integrate it into RAG pipelines that leverage ArangoDB's multi-model capabilities. @arangoai flexible schema and native support for documents, graphs, and vectors make it ideal for applications that need both structured data storage and semantic search. Whether you're indexing documents, managing relational data with embeddings, or building hybrid search experiences, ArangoDB handles it seamlessly. 🐍 pip install arangodb-haystack 🔗 haystack.deepset.ai/integrat…
7
277
A new Document Store has been added to Haystack. Use @supabase to build pgvector vector search capabilities directly into your Haystack pipelines, use Groonga for efficient full-text search, or choose Supabase Storage for seamless file handling. @supabase makes it easy to host your vector database in the cloud without managing infrastructure - just spin up a project, get connection credentials, and you're ready to embed and search. This works well for applications needing scalable vector storage with the familiarity and power of @PostgreSQL 🐍 pip install supabase-haystack 🔗 haystack.deepset.ai/integrat…
1
1
7
728
Extract text, tables, and forms from images and PDFs with Amazon Textract in Haystack. AmazonTextractConverter brings @awscloud's OCR capabilities directly into your document processing pipelines. Just pass an image or single-page PDF and ask natural-language questions to extract structured data automatically. 🐍 pip install amazon-textract-haystack 🔗 haystack.deepset.ai/integrat…
1
8
171
👋 @LiteLLM is now available as a Chat Generator in Haystack. Access 100+ LLM providers from a single interface - OpenAI, Anthropic, Groq, Cohere, and more. Switch models and providers without rewriting code. 🐍 pip install litellm-haystack 🔗 haystack.deepset.ai/integrat…
1
2
11
397
Haystack now publishes a public Model Context Protocol server. Point Claude Code, Cursor, or any MCP-compatible coding agent at it and get real-time access to the latest Haystack docs - no API keys or sign-ups needed. 🖇️ docs.haystack.deepset.ai/doc…
2
3
13
587
Haystack 2.30 is here 📷 Smarter code processing and simpler APIs. Two highlights this time.
1
2
12
443
Vespa is now available as a Document Store in Haystack. Use VespaDocumentStore for hybrid and semantic search with a powerful, production-ready engine, and pair it with VespaEmbeddingRetriever to index and retrieve documents directly in your pipelines. Metadata filtering included. @vespaengine excels at large-scale information retrieval with advanced features like real-time indexing, multi-modal search, and distributed document management - ideal for applications that demand both speed and sophistication. 🐍 pip install vespa-haystack 🔗 haystack.deepset.ai/integrat…
1
10
370
We're launching the Haystack Ambassador Program 🎉 It's for the people already building, teaching, contributing to the Haystack ecosystem and who want to go further. 🧵 Here's what it is, who it's for, and how to apply
6
3
32
1,826
🌐 @perplexity_ai is now integrated with @Haystack_AI! - `PerplexityWebSearch` to fetch fresh web results directly into your pipelines - `PerplexityChatGenerator` to power agent conversations with Perplexity's OpenAI-compatible LLM API - `PerplexityTextEmbedder` and `PerplexityDocumentEmbedder` for semantic embeddings The web search component is especially useful when your agents need real-time information beyond their training data. Just instantiate it with your API key, set `top_k` results, and pass your query. You get back both documents and source links in a single call. Perplexity's API combines language models with live web access, making it a strong choice for building knowledge-intensive applications that need current information. 🐍 pip install perplexity-haystack 🧩 haystack.deepset.ai/integrat…… 📖 docs.perplexity.ai
5
11
2,219
We're so happy that @e2b is now available as a tool integration in Haystack. Give your agents direct access to isolated Linux sandboxes for code execution, file operations, and command-line interactions - all without leaving your pipeline. E2BToolset lets agents write scripts, run code, and explore filesystems inside secure, ephemeral environments. Perfect for coding assistants, data processing tasks, or any agent that needs to execute code safely. 🐍 pip install e2b-haystack 🔗 haystack.deepset.ai/integrat…
2
13
1,049
Use AlloyDBDocumentStore to build RAG applications backed by Google Cloud's PostgreSQL-compatible database, with native vector search and full Haystack pipeline integration. Perfect for enterprises running workloads on @googlecloud who need scalable, managed vector storage alongside relational data. 🐍 pip install alloydb-haystack 🔗 Documentation: haystack.deepset.ai/integrat…
6
212
DoclingServeConverter integrates docling-serve into Haystack pipelines via HTTP, eliminating the need to embed Docling's dependencies in your application. Configure the converter with a docling-serve endpoint, send documents through your pipeline, and receive structured output with Layout and Structure information. This architecture decouples document processing from your main service, enabling independent scaling and reducing deployment overhead. 🐍 pip install docling-serve-haystack 🔗 haystack.deepset.ai/integrat… @LFAIDataFdn
2
7
319
FalkorDB is now available as a Document Store in Haystack. Use FalkorDBDocumentStore to store documents and vectors in a graph format, pair it with FalkorDBEmbeddingRetriever to retrieve documents by semantic similarity, and use FalkorDBCypherRetriever for graph-based queries. This makes it easy to build retrieval pipelines that leverage both vector search and graph relationships. @falkordb is built as an extension of Redis, making it lightweight and efficient for applications that need real-time graph and vector capabilities. 🐍 pip install falkordb-haystack 🔗 haystack.deepset.ai/integrat…
1
3
11
3,748
@brave Search is now available as a web search component in Haystack. Use BraveWebSearch to retrieve relevant documents from the web and pipe them directly into your LLM-powered applications. It integrates seamlessly with Haystack's Pipeline, making it easy to build augmented generation pipelines that leverage real-time web search results. @brave Search combines privacy-respecting search with developer-friendly API design. It's ideal for building agents, question-answering systems, or any application that needs fresh information from the web without relying on a separate search service. 🐍 pip install brave-search-haystack 🔗 haystack.deepset.ai/integrat…
3
98
Asqav is now available as a governance and audit trail integration for Haystack. Use AsqavComponent to sign data flowing through your pipelines with ML-DSA-65 post-quantum signatures and create tamper-evident audit records. Each signed action returns a signature_id linking to its audit trail. Asqav provides governance and compliance capabilities for organizations that need auditable, tamper-proof records of data processing in their AI systems. 🐍 pip install asqav 🔗 haystack.deepset.ai/integrat…
1
4
160
@ChonkieAI is now available as a text-splitting integration in Haystack. Use ChonkieTokenDocumentSplitter and other chunkers from Chonkie's library to intelligently segment your documents before indexing. The integration includes RecursiveChunker, SemanticChunker, and additional chunking strategies to match your workflow. @ChonkieAI's chunkers are designed to preserve semantic boundaries and handle various text types. Whether you need simple token-based splitting or sophisticated semantic awareness, the integration exposes multiple chunking strategies that can be wired into your preprocessing pipelines and retrieval systems. @ChonkieAI has become a popular choice in RAG systems for its ability to balance chunk quality with performance, and now it integrates seamlessly into Haystack's document processing workflows. 🐍 pip install chonkie-haystack 🔗 haystack.deepset.ai/integrat…
6
429
Haystack 2.29 is here 🚀 This release brings hybrid search and flexible LLM messaging to the core, plus async support for cache checking. Many interesting updates this time, but only one highlight: 🔍 Combine Retrievers with MultiRetriever Run multiple text retrievers in parallel and merge their results into a deduplicated list ranked by reciprocal rank fusion. Toggle individual retrievers at runtime using the active_retrievers parameter - skip the embedding retriever for keyword-only queries, for example. 💙 Big thanks to our contributors to this release! 👇 Full release notes below
1
1
5
184
Presidio is now available as a PII detection and anonymization integration in Haystack. Use PresidioEntityExtractor to identify sensitive information like email addresses, phone numbers, names, and credit card numbers in your documents and text. Pair it with PresidioDocumentCleaner or PresidioTextCleaner to automatically redact or mask that data before it flows through your pipelines. This is essential for applications handling user data at scale - financial systems, healthcare platforms, customer support tools, or any RAG pipeline that processes documents with personal information. Configure entity types and confidence thresholds to tune detection precision for your use case. Presidio, built and maintained by @Microsoft, has become the industry standard for PII detection and is now deeply integrated into Haystack workflows. Multi-language support means your data governance works across global datasets.
1
6
255