a voice engineer worries about time-to-first-token. a content engineer worries about cost-per-output. put different contexts at the same table and see where they diverge, converge, and what we can infer.
we're putting together a small technical afternoon for researchers in the space and engineers running inference systems day to day, across different optimisation regimes, latency-obsessed, cost-obsessed, mixed tradeoffs, plus infra people who watch the pattern repeat across all of them.
ten to twelve people, saturday, august 29th with our friends at @lossfunk. we start from different priority metrics and build a shared framework for thinking about inference from wherever the insights lead.
deep in inference and want in? link in the comments.
1
8
1,644


