News, community, and courses for people building AI-powered products.

🌍🌎🌏
🥞🦜 Full Stack LLM Bootcamp 🦜🥞 tl;dr We're releasing our lectures on building LLM-powered apps, for FREE. 🚀 Launch an LLM App in One Hour ✨ Prompt Engineering 🗿 LLM Foundations 🔨 Augmented LLMs 🤷 UX for LUIs 🏎️ LLMOps 🔮 What's Next? 👷 Project Walkthrough Learn more:
18
187
1,029
537,446
The Full Stack retweeted
In the course of running Gradescope, we learned what makes for effective evaluation. It's being able to: • create descriptive rubric items • apply multiple rubric items to the thing being evaluated • and -- crucially -- change rubric item weights on the fly When we set out to evaluate different multiplayer AI platforms, we decided to apply the same principles. We created a set of 18 criteria, such as "In a shared conversation, the agent only uses connectors everyone present has access to." We evaluated each platform (Claude Tag, Dust, QM, Superconductor, and Viktor) on each criterion, recording videos so that you can see for yourself. And we presented the results in an interactive table, where you can change the weights of each item, coming up with a score that represents your team's needs. Curious to hear what you think! multiplayer-ai.com/compare-p…
3
2
13
1,763
The Full Stack retweeted
This is the most interesting recent benchmark result that I've seen: The 100-line mini-swe-agent harness gets better performance out of Opus, GPT, and Gemini than their respective bespoke harnesses. (As measured on the excellent DeepSWE bench). Why would that be true?
10
3
37
8,486
Would you be interested in a course or workshop on ✨Building Software with AI Agents✨???
79% Yes
21% No
66 votes • Final results
3
2
9
3,330
The Full Stack retweeted
What percentage of your Twitter feed (the stuff you actually read, not just scroll past) do you believe is currently written by AI?
55% 0-5%
34% 6-25%
11% >25%
159 votes • Final results
1
2
2
3,930
The Full Stack retweeted
LLM Provider Comparisons 1. @withmartian 2. @ArtificialAnlys 3. @FixieAI
1
5
23
3,968
The Full Stack retweeted
Has anyone done comprehensive testing of gpt-4-vision-preview? I want to know stuff like the minimum text size it can read, the radius of the smallest circle it can locate in an image, the number of circles it can count, etc. Could be an automated benchmark for other models too
3
1
13
3,508
The Full Stack retweeted
Which set of statements do you agree with? 1. AGI is as much or more of a risk to human flourishing as nuclear weapons 2. I have a good idea for what should be done about that
62% {}
28% {1}
10% {1,2}
108 votes • Final results
1
1
4
3,088
The Full Stack retweeted
Has anyone had good experiences with GPT-powered code generation for complete web app features? As in, you describe what should exist, and GPT actually provides the source of all the necessary files and where they should go. Ideally in the context of Ruby on Rails.
9
2
9
5,091
The Full Stack retweeted
Let's say that a US-based research company has developed an AGI model that was able to use the browser, pass captchas, hire people on Upwork, and lie about its intentions. What should they do after observing this?
38% Open-source the weights
27% Be loud and publish paper
19% Be quiet and alert gov
16% Nothing; keep researching
586 votes • Final results
6
7
19
10,364
The Full Stack retweeted
We bring in @full_stack_dl, a venerable boot camp crew that pioneered technical deep dives into deep learning where people fly in from around the world. 🥞 Their #LLM Bootcamp in the spring was sold out and this is your chance to attend the ➡️ version. 👉 scale.bythebay.io/register
1
4
2,370
We're live to talk about production AI, LLMs, open source, and more! piped.video/watch?v=aN3OxHj2…
6
37
4,464
We're hosting a livestream with @ScaleByTheBay, this coming Monday at 1:30 pm PST. Come join us on your YouTube channel to talk about LLMs in production and more. piped.video/@The_Full_Stack)
1
6
27
3,486
We're also about 3 weeks away from our latest LLM bootcamp. @karpathy called the last version "high-quality tokens". Register soon if you want to make sure you get a spot! The bootcamp is in Oakland on November 13. You can register here: scale.bythebay.io/llm-worksh….
1
5
1,836
The Full Stack retweeted
Solutions from replies: - @OpenPipeAI looks exactly right openpipe.ai - @PortkeyAI launching feature soon - @analyticsaurabh building his own I currently use @helicone_ai, any plans from them?
Is there a service I can use to pipe my GPT-4 calls through, and it automatically finetunes GPT-3.5 (or whatever) on all of them, and lets me know when it's up to par?
4
7
33
10,617
The Full Stack retweeted
Wow - don't miss this!
This is sadly true! If you want the latest version, come join us in November for our in-person workshop with @ScaleByTheBay scale.bythebay.io/llm-worksh…
1
6
3,838
The Full Stack retweeted
Is there a service I can use to pipe my GPT-4 calls through, and it automatically finetunes GPT-3.5 (or whatever) on all of them, and lets me know when it's up to par?
8
4
32
15,485
This is sadly true! If you want the latest version, come join us in November for our in-person workshop with @ScaleByTheBay scale.bythebay.io/llm-worksh…
Just realized that even the best and most up-to-date #LLMbootcamp from @full_stack_dl is partially outdated! The field is rushing! #LLM #fullstack
4
19
8,400
The Full Stack retweeted
It feels like something to be you. Do you think it feels like something to be GPT-4?
17% Yes
68% No
15% Show results
213 votes • Final results
5
1
4
4,138
The Full Stack retweeted
By what year will there be an AI that is more capable than most humans in most domains of digital work (e.g. you can tell it to do anything you currently hire a white collar professional to do, and it does the job better than the median human)?
30% 2025
45% ~2033
17% ≥2040, but in my lifetime
8% not in my lifetime
152 votes • Final results
3
1
2
3,104