Humanloop is the LLM evals platform for enterprises. Trusted by Gusto, Vanta and Duolingo to ship reliable AI products.

SF and London
Based in United States
Filter
Exclude
Time range
-
Minimum likes
We're thrilled to announce that the Humanloop team is joining @AnthropicAI! Our mission has always been to enable the rapid and safe adoption of AI. Now, as AI progress accelerates, we think Anthropic is the ideal home to continue this work.
27
21
459
243,541
MCP is rapidly becoming the universal adapter for AI. Since its release in November, developers and teams have raced to adopt the standard, giving agents the tools they need to interface with the real world, from APIs to internal systems. Our latest explainer breaks down MCP: what it is, how it works, and how to get started. 🔗 Read here: humanloop.com/blog/mcp
12
1
14
2,189
🗓️ Wednesday March 12th at 10:00 PT Our CEO @RazRazcle will be speaking at the MLOps Community’s 'AI in Production 2025' about Eval-Driven AI Development. What to expect: • Learn how top AI teams use evaluation-driven development to guide model improvements and avoid common pitfalls. • Discover how to leverage code-based, LLM-as-judge, and human evaluators to optimize LLM performance. • Gain insights from Brianna Connelly, VP of Data Science at @filevine, on how their AI team uses evals on Humanloop to refine AI applications and RAG systems. Register to take part virtually (link below)
2
5
1,258
📍PMs in AI Meetup, London 🇬🇧 Yesterday we held a Meetup in the UCL Centre for Artificial Intelligence for product managers working on AI agents and applications. Huges thanks to all who turned up (it was a full house!) and to our speakers: • @samstphenson (Founder, @meetgranola) - who advised on making your 1 AI feature extremely effective before trying to add any more. • @Albertorizzoli (Co-founder, @V7labs) - who said to listen to user problems, not their proposed solutions (this is more true than ever with AI). • @RazRazcle (Co-founder, @humanloop) - advised to bring domain experts into the prompt engineering and evaluation process as early as possible to drive differentiated and effective AI performance. The London AI community is next level 🚀 What should be the theme of our next meetup? 👀
3
1
25
4,212
How do you take your AI product from good to great to game-changing? Next Tuesday (Feb 18th) we’re hosting a Product Managers in AI Meetup in Bloomsbury, London 🇬🇧 Join us for a panel on "How to Build AI Products That Delight Users" with guest speakers: • @samstphenson, Founder, @meetgranola • @Albertorizzoli, Co-founder, @V7Labs • @RazRazcle, Co-founder, @humanloop Food and drinks will be provided. Limited availability — register here: lu.ma/gwojmdql
1
4
872
Today we’re introducing Templates - a library of Prompts, Evaluators, and Datasets, designed to accelerate time to value when developing and evaluating AI applications. One of the biggest challenges in testing AI applications and agents is accessing the right datasets and evaluators. So we’ve collaborated with @huggingface to make this easier. With Templates, the best and most popular golden datasets on Hugging Face are instantly accessible in Humanloop, alongside our fully customizable pre-set evaluators, to help you streamline LLM evaluations. No more starting from scratch - easily test your prompts and agents for jailbreak vulnerabilities, PII leaks, text-to-SQL accuracy, domain-specific reasoning, and lots more — powered by @huggingface Datasets and @humanloop Evals. Templates are live now! (Link below to learn more).
1
1
8
1,112
When do you know it's time to try fine-tuning instead of prompt engineering? Our CEO @RazRazcle is on Data Radicals with @satyx this week to discuss: 🔹 How fine-tuning tends to be an optimization step, which comes once you've pushed the limits of prompt engineering 🔹 Why collaboration with domain experts in the AI product development cycle is key to driving successful outcomes 🔹 How software engineering is changing in the age of AI And lots more! Watch the full episode here: alation.com/podcast/episodes…
1
2
783
o3-mini now available in the prompt editor — with streaming!
OpenAI o3-mini is now available in ChatGPT and the API. Pro users will have unlimited access to o3-mini and Plus & Team users will have triple the rate limits (vs o1-mini). Free users can try o3-mini in ChatGPT by selecting the Reason button under the message composer.
3
593
Replying to @RuiCarrilho5
May your timeline be blessed with fewer regrettable minutes.
1
10
2,157
only if you want to add a lot of noise to your timeline!
2
26
6,170
Replying to @0PointAlchemy
patience, sir.
9
4,704
Interact to clean your timeline: Transformers LLMs Prompt engineering CoT Constitutional AI Ray Kurzweil predictions TPU pods Attention is all you need noam shazeer Scaling laws Colab Pro ooms GPT wrappers Model distillation AGI timelines p doom NVDA VRT TSMC open weights unhobblings infinite context stargate
332
246
5,196
190,271
“You wouldn’t build a $100m software product without unit tests. Then how can you think of building a $100m AI product without evals?” - Noam Rubin, AI Platform team at @TrustVanta Noam joined us in SF to speak about how his team have used Humanloop to build some of the most compelling AI products on the market. Noam spoke about the differences between traditional software development and building with AI. "Most engineers haven't built with stochastic software before, and so teaching them about how to use evals and datasets in iterative deployment has been key" Noam's team use Humanloop to run evaluations, which is now part of their CI/CD workflow. "We don't ship a prompt change now unless it has an eval report from Humanloop. Its literally in the PR" Thanks for coming by Noam! We're stoked to be supporting you.
3
752
2024 was about shipping AI products that work. In our mission to make this easier, this year we had: • 50 product releases • 50 new models supported • 300 production deployments Resulting in thousands of new AI products being deployed on Humanloop 🧵
1
1
7
741
Introducing the AI Engineer Pack! Get $50+ in credits from each of the leading AI developer tools including @humanloop, @elevenlabs and more Whether you’re building a new AI product at work or launching a side project, the AI Engineer Pack has everything you need to build with AI
1
1
5
507
“I'm convinced the vast majority of companies leveraging generative AI today are operating in the dark” - Brianna Connelly, VP of Data Science at @filevine Brianna joined us in SF to talk about Filevine’s journey to becoming the legal tech stack supercharged by AI. When building out their first AI feature, legal domain experts would manually prototype prompts before handing them off to engineers to go live into production - leaving them with no visibility into performance or ability to make changes to prompts once the product went live. “Our prompt management and evaluation process was extremely manual and time-consuming, done entirely on spreadsheets. This created a significant bottleneck that slowed down our product roadmap and prevented us from adopting new models.” Brianna came to Humanloop to solve this. By unifying her team’s AI workflows around prompt engineering, evaluation and observability on Humanloop, Filevine drastically improved the performance and reliability of their AI product. Since then, they’ve shipped 6 new AI products and are saving over 16 hours per week on evaluation & prompt management! Thanks so much for support Brianna! We’re incredibly proud to be enabling your team to ship and scale AI with confidence.
3
424
Photos are live from our SF Launch event! 🇺🇸 Huge thanks to everyone who came down and helped make it such a special evening. Stay tuned for what's next!
1
3
445