Humanloop is the LLM evals platform for enterprises. Trusted by Gusto, Vanta and Duolingo to ship reliable AI products.

SF and London
We're thrilled to announce that the Humanloop team is joining @AnthropicAI! Our mission has always been to enable the rapid and safe adoption of AI. Now, as AI progress accelerates, we think Anthropic is the ideal home to continue this work.
27
21
459
243,539
MCP is rapidly becoming the universal adapter for AI. Since its release in November, developers and teams have raced to adopt the standard, giving agents the tools they need to interface with the real world, from APIs to internal systems. Our latest explainer breaks down MCP: what it is, how it works, and how to get started. 🔗 Read here: humanloop.com/blog/mcp
12
1
14
2,187
🗓️ Wednesday March 12th at 10:00 PT Our CEO @RazRazcle will be speaking at the MLOps Community’s 'AI in Production 2025' about Eval-Driven AI Development. What to expect: • Learn how top AI teams use evaluation-driven development to guide model improvements and avoid common pitfalls. • Discover how to leverage code-based, LLM-as-judge, and human evaluators to optimize LLM performance. • Gain insights from Brianna Connelly, VP of Data Science at @filevine, on how their AI team uses evals on Humanloop to refine AI applications and RAG systems. Register to take part virtually (link below)
2
5
1,256
📍PMs in AI Meetup, London 🇬🇧 Yesterday we held a Meetup in the UCL Centre for Artificial Intelligence for product managers working on AI agents and applications. Huges thanks to all who turned up (it was a full house!) and to our speakers: • @samstphenson (Founder, @meetgranola) - who advised on making your 1 AI feature extremely effective before trying to add any more. • @Albertorizzoli (Co-founder, @V7labs) - who said to listen to user problems, not their proposed solutions (this is more true than ever with AI). • @RazRazcle (Co-founder, @humanloop) - advised to bring domain experts into the prompt engineering and evaluation process as early as possible to drive differentiated and effective AI performance. The London AI community is next level 🚀 What should be the theme of our next meetup? 👀
3
1
25
4,211
Very excited to share that Humanloop is named as one of Emerging Leaders in the Emerging Market Quadrant for Generative AI Engineering in the 2025 Gartner® Innovation Guide for Generative AI Technologies. Building reliable AI products doesn't have to be a guessing game - you need an eval-driven workflow and the right collaboration between engineers, product and domain experts. That’s how teams ship AI that works. Companies like Gusto, Vanta & Duolingo use Humanloop to build great AI products, we love helping companies navigate this new paradigm - reach out to ask questions or book a demo with us humanloop.com Read more in Gartner 2025 February edition of the Innovation Guide for Generative AI Technologies
1
681
🚀 We’re thrilled to be part of ProductCon on Feb 19, 2025, in London! ProductCon brings together the brightest minds in tech to share best practices for building world-class products. 📍 Stop by our booth to see how PMs use Humanloop to build differentiated AI products that perform reliably at scale. 🔗 Grab your ticket here: lnkd.in/e_dEnQqX
1
711
How do you take your AI product from good to great to game-changing? Next Tuesday (Feb 18th) we’re hosting a Product Managers in AI Meetup in Bloomsbury, London 🇬🇧 Join us for a panel on "How to Build AI Products That Delight Users" with guest speakers: • @samstphenson, Founder, @meetgranola • @Albertorizzoli, Co-founder, @V7Labs • @RazRazcle, Co-founder, @humanloop Food and drinks will be provided. Limited availability — register here: lu.ma/gwojmdql
1
4
872
Today we’re introducing Templates - a library of Prompts, Evaluators, and Datasets, designed to accelerate time to value when developing and evaluating AI applications. One of the biggest challenges in testing AI applications and agents is accessing the right datasets and evaluators. So we’ve collaborated with @huggingface to make this easier. With Templates, the best and most popular golden datasets on Hugging Face are instantly accessible in Humanloop, alongside our fully customizable pre-set evaluators, to help you streamline LLM evaluations. No more starting from scratch - easily test your prompts and agents for jailbreak vulnerabilities, PII leaks, text-to-SQL accuracy, domain-specific reasoning, and lots more — powered by @huggingface Datasets and @humanloop Evals. Templates are live now! (Link below to learn more).
1
1
8
1,112
When do you know it's time to try fine-tuning instead of prompt engineering? Our CEO @RazRazcle is on Data Radicals with @satyx this week to discuss: 🔹 How fine-tuning tends to be an optimization step, which comes once you've pushed the limits of prompt engineering 🔹 Why collaboration with domain experts in the AI product development cycle is key to driving successful outcomes 🔹 How software engineering is changing in the age of AI And lots more! Watch the full episode here: alation.com/podcast/episodes…
1
2
783
Release notes 01/31/2025 New models: • o3-mini is now available in Humanloop! It has 200,000 token context length, with 100,000 output tokens and it can show superior performance to o1, but for 9x cheaper and 4x faster. • DeepSeek V3 and R1 (non-distilled) via DeepSeek API. We also added DeepSeek-R1-Distill-Llama-70B via Groq. Evals: • Quickly compare performance of various prompts/models using aggregate eval stats in the UI (see below). • Better filtering across evals for errors as well as specific judgements and models. Read more (link below)
1
1
706
🗓️ Tomorrow! (Thurs 15th) Our CEO Raza Habib will be providing a workshop on Eval-driven AI Product Development at the AI Builders Summit. You'll learn about: • Designing evaluators for building reliable AI applications, including RAG systems and agents. • Using code-based, LLM-as-judge, and human evaluators to optimize performance. • Incorporating evaluation-driven development to avoid pitfalls and improve iterative development processes. Register to take part virtually (link below)
2
1
1,033
Release notes 01/15/2025: • New UI layout: refined sidebars (supporting drag & drop), tabs, and side panels that create a more cohesive experience developing and evaluating AI apps. • Full trace visibility when evaluating Agents • Enhanced editor for creating AI evaluators (supporting function calling by default) • Quick eval runs: multi-select logs and trigger an evaluation run • Latest models: Llama 3.3 70B with blistering speed on @GroqInc and @GeminiApp’s Experimental Models (gemini-2.0-flash-exp and gemini-exp-1206) Also — decorators for quickly integrating Humanloop via our Python and Typescript SDKs, support for user-defined log ids and much more (link below). Happy 2025!
1
1
855
“We strike a balance between automation and human expertise, ensuring people stay at the heart of creative decisions” - Nikesh Hotchandani, AI Product Owner at Tag Nikesh joined us in London to talk about how he’s deploying AI in a responsible and scalable manner for one of the world’s largest marketing production agencies. He spoke about how LLMs are transforming marketing but in order for his team to scale AI production, they must first ensure model output is aligned with Tag’s global standards. To solve this, Nikesh and his team made evaluations a core part of their workflow - and now they leverage Humanloop to ensure all of their prompts align with company guidelines and perform responsibly We’re proud to be supporting Nikesh’s team at Tag! Thanks for coming by.
1
599
“You wouldn’t build a $100m software product without unit tests. Then how can you think of building a $100m AI product without evals?” - Noam Rubin, AI Platform team at @TrustVanta Noam joined us in SF to speak about how his team have used Humanloop to build some of the most compelling AI products on the market. Noam spoke about the differences between traditional software development and building with AI. "Most engineers haven't built with stochastic software before, and so teaching them about how to use evals and datasets in iterative deployment has been key" Noam's team use Humanloop to run evaluations, which is now part of their CI/CD workflow. "We don't ship a prompt change now unless it has an eval report from Humanloop. Its literally in the PR" Thanks for coming by Noam! We're stoked to be supporting you.
3
752
2024 was about shipping AI products that work. In our mission to make this easier, this year we had: • 50 product releases • 50 new models supported • 300 production deployments Resulting in thousands of new AI products being deployed on Humanloop 🧵
1
1
7
741
Introducing the AI Engineer Pack! Get $50+ in credits from each of the leading AI developer tools including @humanloop, @elevenlabs and more Whether you’re building a new AI product at work or launching a side project, the AI Engineer Pack has everything you need to build with AI
1
1
5
507
“I'm convinced the vast majority of companies leveraging generative AI today are operating in the dark” - Brianna Connelly, VP of Data Science at @filevine Brianna joined us in SF to talk about Filevine’s journey to becoming the legal tech stack supercharged by AI. When building out their first AI feature, legal domain experts would manually prototype prompts before handing them off to engineers to go live into production - leaving them with no visibility into performance or ability to make changes to prompts once the product went live. “Our prompt management and evaluation process was extremely manual and time-consuming, done entirely on spreadsheets. This created a significant bottleneck that slowed down our product roadmap and prevented us from adopting new models.” Brianna came to Humanloop to solve this. By unifying her team’s AI workflows around prompt engineering, evaluation and observability on Humanloop, Filevine drastically improved the performance and reliability of their AI product. Since then, they’ve shipped 6 new AI products and are saving over 16 hours per week on evaluation & prompt management! Thanks so much for support Brianna! We’re incredibly proud to be enabling your team to ship and scale AI with confidence.
3
424
Kicking off our GA launch in London 🇬🇧 The AI community here is incredible.
2
470
Photos are live from our SF Launch event! 🇺🇸 Huge thanks to everyone who came down and helped make it such a special evening. Stay tuned for what's next!
1
3
445
We’re thrilled to share that Humanloop has been named an Emerging Leader in Gartner Emerging Market Quadrant for Generative AI Engineering. The AI/LLM landscape is rapidly evolving and we love helping companies navigate this new paradigm — reach out to ask questions or book a demo with us humanloop.com Learn more about Gartner’s report in the 2024 November edition of the Innovation Guide for Generative AI Technologies (subscription required)
2
517
“You wouldn’t build a $100m software product without unit tests then how can you think of building a $100m ai product without evals” Noam Rubin from Vanta
1
4
456
We're most proud of what our customers have been able to achieve: @filevine lawyers/engineers are collaborating to reimagine legal work with AI @GustoHQ are building agents, reinventing internal processes and product @DixaApp is transforming customer support at scale
1
6
631
🔹Look at their data, constantly 🔹Iterate against custom evaluators 🔹Let domain experts drive development We built Humanloop to productise these best practices so your entire team can adopt eval-driven development.
1
4
744
Humanloop is now generally available! After 2 years of working closely with early customers, we're opening access to our full evals platform. 🧵 Here's what we've learned and how we can help you build great AI products:
17
14
112
38,780
How can you make sure your AI application isn’t going to do something you don’t want it to do? LLM guardrails are designed to mitigate output risks by aligning model behaviour with intended use. Our latest guide covers LLM guardrails. You can expect to learn: • The different types of guardrails enterprises commonly use • Which ones are most important and why • Best practices their for implementation (link below)
1
3
419
When scaling your LLM application, how do you maintain consistent high quality performance? Our latest guide covers LLM monitoring, the ongoing process of observing and analyzing the performance of an LLM app once it's deployed. You'll learn: • How LLM monitoring works and why it's important • Best practices and most common metrics used when setting it up • How monitoring is different to evaluations • Common pitfalls to be aware of (link in replies)
1
1
5
446
That’s a wrap for our company offsite! We spent a week in New York bonding, catching up with customers and working on their most frequent asks. We also had time to have a camp fire, compete in VR and do a scavenger hunt around the city 🗽 Exciting news coming soon…
3
298
At #devday this week @OpenAI released prompt caching - which reduces input token costs by 50% and latency by 80% 🤯 Not too long ago, @AnthropicAI released its version of prompt caching, which can save up to 90% on input tokens and reduce latency by 85% 🚀 This is a huge deal. Not only because they will make AI applications faster and cheaper, but because they require you to re-think how you engineer your prompts. Depending on which model provider you use, prompt caching varies in terms of cost, control and complexity. To understand the full picture, read our guide (link below)
1
2
6
659
How does a traditional machine learning team transition into generative AI? @DixaApp was asking themselves this question in late 2022, when they realized they could achieve a lot more with less by leveraging LLMs. Using Humanloop, Dixa applied their machine learning expertise to generative AI app development, enabling them to evaluate and track key LLM performance metrics Now Dixa releases AI products 3x faster and achieves 95-100% accuracy rates across all LLM features. Read more below ⬇️
1
1
1,336
Reduce token costs by 90% and latency by 85% in your RAG application? 🤯 @AnthropicAI recently released prompt caching, a technique for making LLMs more efficient over long-context tasks and it’s super impressive. Check out our latest explainer to learn: 🔸 How prompt caching works in Anthropic and framework’s like CacheGPT 🔸 Where prompt caching can be most useful (and not so useful) 🔸 The best applications of prompt caching today Link below ⬇️
1
3
563
You can now use @OpenAI's o1 on Humanloop! o1 boasts incredible reasoning capabilities and can solve harder problems in science, coding, and math than it's predecessor, GPT-4o. Bring your OpenAI API key to test, evaluate and deploy o1 now: humanloop.com
1
4
526
How much faster could you develop AI features if you had 100% visibility into LLM performance? @DixaApp is one the fastest growing customer service platforms in the world, processing millions of conversations each month. To develop generative AI features that are reliable at scale, Daniele Alfarone and his team of experienced product and machine learning engineers knew they needed a unified platform for managing, evaluating, and monitoring LLMs. By leveraging Humanloop, they’ve not only achieved 95-100% performance accuracy on all AI features but also accelerated development timelines by 300%. Read the full story below ⬇️
1
1
749
⭐Customer Spotlight: How @filevine revolutionized the legal tech AI landscape… In 12 months: - 6 new AI products - 2x AI revenue - 16 hours saved weekly - Inc. 5000 recognition How? Single source of truth for GenAI management with Humanloop. Full story: hubs.ly/Q02JP4gP0 #BuiltWithHumanloop
1
1
5
1,640
Want to break into AI engineering? In this week's episode of the High Agency Podcast @jxnlco and @RazRazcle share actionable advice for building with #LLMs.
1
2
6
998
New research paper from our CTO Peter Hayes presented at #ICML2024 this week in Vienna titled “Active Preference Learning for Large Language Models.” If you’re interested in going deeper on #LLMops use cases for fine-tuning and RLHF, you can check out the abstract and full paper here: hubs.ly/Q02HGycl0 We'll be onsite through EOD Wednesday. DM if you’d like to meet-up 🇦🇹
1
4
475
🚨 Introducing GPT4o-mini 🚨 New model launched by @OpenAI with: • 128k context window • 60 cents per 1M output tokens • Support for text and vision API Bring your API key to Humanloop now to test, evaluate and deploy!
1
3
1,196
This week on the High Agency Podcast, Humanloop co-founder & CEO @RazRazcle sits down with @OfficialLoganK to discuss the future trajectory of AI. ... and his predictions might surprise you. Logan led developer relations at @OpenAI before leading product on the @googleaistudio. He's been closer than anyone to developers building with #LLMs and has seen behind the curtain at two frontier labs. Topics discussed include: 🔸 What it was like joining OpenAI the day ChatGPT hit 1 million users 🔸 What you might expect from GPT5 🔸 Google's latest innovations and the battle with OpenAI 🔸 How can you stay ahead and achieve real ROI 🔸 Logan's insights into the form factor of AI and what will replace chatbots Check out the full episode wherever you listen to podcasts: → YouTube hubs.ly/Q02G_NXr0 → Spotify hubs.ly/Q02G_Q5Q0 → Apple hubs.ly/Q02G_PfH0
1
3
535
"There are advantages to being early but you do not get the right to win just by being early" @Swyx is the founder of the AI Engineer World's Fair and author of "Rise of the AI Engineer". Watch his full conversation with @RazRazcle here: hubs.ly/Q02DLCTM0
1,138
What's the difference between an AI engineer and an ML engineer? "The AI engineer looks after the 0-1 phase whereas the ML Engineer takes on the 1-N phase" @swyx penned the viral essay 'Rise of the AI Engineer' (latent.space/p/ai-engineer) and is the founder of the @aiDotEngineer World's Fair. He joined Raza on the High Agency podcast this week to discuss what an AI engineer really is. Watch: hubs.ly/Q02DLDzj0
2
5
26
7,095
Are you attending the @aiDotEngineer World's Fair this week in SF? If so, you don’t want to miss this session… Today our CEO Raza Habib will be sharing practical insights from AI leaders from the #HighAgencyPodcast as well as Humanloop customers like @duolingo, @GustoHQ, @TrustVanta, @filevine, @ironclad_inc, @sourcegraph, and many more. These trailblazers have successfully implemented #LLMs in production, and we’ll be bringing actionable lessons and best practices for how to get #GenAI apps from prototype to production. A glimpse at what we’ll be covering today at the AI Leadership track at 12:15 p.m. PT: 🔹 The essential skills your team needs to thrive 🔹 Tips for running #RAG in production 🔹 How to choose the right evaluations 🔹What it takes to succeed with AI agents Join us to gain valuable knowledge and actionable strategies to leverage LLMs effectively in your organization. Link to session details below.👇 See you there!
1
2
514
One consistent piece of advice that has comes on almost every episode of the High Agency Podcast... • Start building today❗ @sourcegraph is one of the world's most popular AI coding tools. Their CTO and Cofounder @beyang shares his thoughts on this ⬇️
2
1
10
554
Claude 3.5 Sonnet is now in Humanloop! — 2x the speed — 1/5th the cost — yet smarter than Claude 3 Opus with full support for images, variables and multiple tool calls👇 Start building incredible applications with the latest intelligence from @AnthropicAI in the best interface for developing and evaluating your prompts
1
1
8
4,506
"There's a huge opportunity for product experts to go and test out a lot of the ideas that they previously couldn't" @sourcegraph CTO @beyang on how lowering the barrier to software development has created a wave of opportunity for domain experts to build new things.
1
1
665
Will agents be fully autonomous or will some degree of human interaction remain necessary/preferable? Bryan Bischof weighs in on where he thinks this is going ⬇️
1
1
5
344
📣 Product Update: Introducing Evaluation Reports! Evaluation reports allow you to create more insightful summaries of your evaluations data on Humanloop. This includes: • Summary statistics with percentiles and standard deviations. • Graphs with spider, box and bar charts for different types of Evaluators. • Clear comparisons of different Prompt/Tool versions. You have the flexibility to extend or duplicate existing reports, as well as to reuse or generate logs for creating new reports 🌟
1
1
1
609
One big lesson from AI engineers who come from data science... ➡️ Spend more time looking at your production data "You will answer most of the hard (AI) product questions by looking at the data. The data is going to generated as you go and it's going to be generated by the users and the agent" - @BEBischof
1
1
3
357