News, community, and courses for people building AI-powered products.

🌍🌎🌏
Filter
Exclude
Time range
-
Minimum likes
The Full Stack retweeted
In the course of running Gradescope, we learned what makes for effective evaluation. It's being able to: • create descriptive rubric items • apply multiple rubric items to the thing being evaluated • and -- crucially -- change rubric item weights on the fly When we set out to evaluate different multiplayer AI platforms, we decided to apply the same principles. We created a set of 18 criteria, such as "In a shared conversation, the agent only uses connectors everyone present has access to." We evaluated each platform (Claude Tag, Dust, QM, Superconductor, and Viktor) on each criterion, recording videos so that you can see for yourself. And we presented the results in an interactive table, where you can change the weights of each item, coming up with a score that represents your team's needs. Curious to hear what you think! multiplayer-ai.com/compare-p…
3
2
13
1,763
Would you be interested in a course or workshop on ✨Building Software with AI Agents✨???
79% Yes
21% No
66 votes • Final results
3
2
9
3,330
We're live to talk about production AI, LLMs, open source, and more! piped.video/watch?v=aN3OxHj2…
6
37
4,463
We're also about 3 weeks away from our latest LLM bootcamp. @karpathy called the last version "high-quality tokens". Register soon if you want to make sure you get a spot! The bootcamp is in Oakland on November 13. You can register here: scale.bythebay.io/llm-worksh….
1
5
1,837
We're hosting a livestream with @ScaleByTheBay, this coming Monday at 1:30 pm PST. Come join us on your YouTube channel to talk about LLMs in production and more. piped.video/@The_Full_Stack)
1
6
27
3,486
This is sadly true! If you want the latest version, come join us in November for our in-person workshop with @ScaleByTheBay scale.bythebay.io/llm-worksh…
Just realized that even the best and most up-to-date #LLMbootcamp from @full_stack_dl is partially outdated! The field is rushing! #LLM #fullstack
4
19
8,399
If you can't make it in-person, check out our materials online instead! Our last LLM bootcamp (as well as our deep learning course) are available for free on our website. fullstackdeeplearning.com/
1
16
1,376
You can register for the course here: scale.bythebay.io/register We hope to see you there!
1
2
1,471
The bootcamp will cover topics like: 🚆 Choosing the right model ✍️ Writing good prompts 🏋️ Deciding when to fine-tune 🏗️ Augmenting LLMs with your own data 🔨 Picking tools and infrastructure By the end, you'll be prepared to start taking on projects of your own with LLMs.
1
4
572
We're excited to partner with our friends at @ScaleByTheBay to bring you this full-day in-person course in Oakland on November 13. You can sign up here: scale.bythebay.io/register We'll teach you how to build LLM-powered applications, from scratch, for production.
1
1
5
971
🥞🦜 New LLM Bootcamp Announcement 🦜🥞 In 2023, the AI world speedran through models, architectures (e.g., RAG), and frameworks (e.g., @langchain). After a year of hype, what's *actually* working? This November, we'll show you, in our latest class on building prod LLM apps
4
18
102
18,789
Come watch some livecoding with LLMs! On our YT channel August 11, 8 am Pacific, live and then forever after: piped.video/live/GaV_M2l4-N0…
.@jeremyphoward put out a delightful tutorial this week on getting started with LLMs for a science QA Kaggle competition unlike many other intros it emphasizes exactly the right thing: understand the data first, then the model, in the context of the task kaggle.com/code/jhoward/gett…
20
85
25,034
Yep, we're interested in RetNet too! But the (literal) million-dollar question with Transformer alternatives is whether they will scale, and the largest RetNet training run has ~2x fewer params and ~3x fewer params than the largest RWKV run.
Replying to @full_stack_dl
(technically, that figure is not for RWKV, but for RetNet, a similar "vectorizable RNN" architecture. it came out while we were writing this tutorial and matches Transformers up to 7B params, 100B tokens. we're watching with interest!) arxiv.org/abs/2307.08621
2
621
Coming soon! RetNet has some really great ideas in it, and we're looking forward to checking out the code.
1
177
We're hoping this walkthrough makes the reasons to be excited clearer to the rest of the community and attracts more work on this model. Big thanks to @BlinkDL_AI, @AiEleuther, and team for pushing the envelope on efficient language modeling! fullstackdeeplearning.com/bl…
3
12
1,801
> What is this blog post? We've written up the RWKV inference code notebook style, with the goal of clarity. That means we: - motivate the architecture step-by-step - use type/shape hints and other documentation - write code to be understood quickly, not run quickly
2
14
4,455
By a bit of napkin math, if RWKV stays on the scaling curve for 7B parameters + 2T tokens you could run a model with the same quality as LLaMA 2's smallest on streaming data on a mobile Jetson Nano GPU, with room to spare for computer vision 👀
1
10
1,000
For RWKV, that memory footprint is small, orders of magnitude smaller than the weights. The 14B parameter RWKV can run on 3GB VRAM -- for _any_ sequence length!
1
1
8
965
(technically, that figure is not for RWKV, but for RetNet, a similar "vectorizable RNN" architecture. it came out while we were writing this tutorial and matches Transformers up to 7B params, 100B tokens. we're watching with interest!) arxiv.org/abs/2307.08621
1
11
1,644
> Why care? At inference time, RNNs use the same amount of memory, no matter how long of a sequence they're operating on. Transformer memory usage scales linearly.
1
9
1,009