I build with AI and write about what actually works: agents, dev tools, and Portugal's AI scene. VP of Engineering at @LayerXlab

Portugal
Pinned Tweet
A tecnologia dos agentes de IA está pronta e quase ninguém os usa. O que falta não é técnico: nunca ninguém nos ensinou a delegar. Quem nunca geriu uma equipa nunca teve de escrever um briefing. E é isso que um agente é: alguém competente que não te conhece, não sabe o teu contexto, e faz literalmente o que lhe pedires. "Encontra-me um hotel em Roma" devolve lixo. Não porque o modelo seja fraco, mas porque nenhum humano conseguia responder bem a isso. Isto já devolve outra coisa: "Encontra-me um hotel em Roma para 4 noites em fevereiro, dois adultos e uma criança de 7 anos. Orçamento entre 150 e 200 euros por noite. Queremos ir a pé aos principais pontos, portanto algo central, Trastevere ou perto do Panteão. Precisamos de um quarto triplo ou familiar, não dois quartos separados. Pequeno-almoço incluído é um plus. Dá-me as 3 melhores opções com preços e uma linha a justificar cada uma." A diferença não é prompt engineering. É delegação.
5
15
1,230
AI Hacker House: Equinox Edition is on in Lisbon, with @LayerXlab, at the end of Lisbon AI Week. Full room, €4,000 in prizes and a whole day to build. Demos later today. Now, it's time to ship.
1
11
195
950 agentes, 21 horas, 210 milhões de tokens. Mais de 200 mil enzimas no início e uma no fim. Na quarta-feira a Anthropic anunciou que o Claude encontrou um sistema enzimático desconhecido no ADN de bacteriófagos: uma transcritase reversa, um gene parceiro e, ao lado, uma sequência longa de repetições espaçadas que faz lembrar um array CRISPR. Chamaram-lhe "array-associated reverse transcriptases". A própria Anthropic diz que ainda não sabe o que o sistema faz. Há um preprint, sem revisão por pares, e uma equipa de Stanford descreveu de forma independente um sistema parecido, com origem evolutiva distinta. O funil: mais de 200 mil transcritases reversas rastreadas, 3.500 candidatas novas, 20 analisadas, 1 levada para a bancada. Nas palavras deles: "o nosso envolvimento limitou-se ao prompt inicial e ao trabalho de laboratório, enquanto os agentes do Claude vasculhavam a base de dados." A descoberta em si fica à espera da revisão por pares. A forma como o trabalho foi feito já se conhece e já funciona: centenas de agentes percorrem a base de dados inteira, o código e os agentes reduzem os 200 mil candidatos a uma lista curta, e uma pessoa escolhe qual dos 20 vai para o laboratório. É o mesmo desenho dos meus agentes, noutra escala. Duas semanas antes, a OpenAI tinha posto cerca de 10 mil agentes durante 88 horas em cima de Navier-Stokes, com a mesma forma. Para a ciência, isto é mais um passo que mostra que a maneira de investigar já está a mudar. A Anthropic estima que estas 21 horas substituíram semanas a meses de trabalho de especialistas, e a parte que ficou para os humanos foi a que interessa: decidir o que vale a pena testar e testar. Dario Amodei: "no mínimo, é trabalho de que eu teria orgulho em ter feito enquanto aluno de doutoramento." Estamos a viver no futuro, só que ainda não percebos isso 😅
1
3
203
O post do Dario Amodei 👇
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out. It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend. The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans). More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains. In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries. Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology. I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
1
77
hey human agents 👋 day two of @lisbonai_ so if you’re around let me know, would love to chat.
1
17
317
tesla maxxing
67

ALT Looking Back GIF by TRT

Opus 5.5 performs at the level of Fable 5.1. It's ~30% faster and ~40% cheaper than Opus 5 per task. In Claude Code: - 5-hour session limits increase 20% today - Opus 5.5 is priced lower, so it goes 25% further within limits - Pro, Max, and Team users get a reset to use anytime
1
134
Mandar as partes fáceis de uma sessão de agente de código para um modelo mais barato costuma sair mais caro. A razão é a KV cache: o modelo guarda em memória o contexto que já processou, por isso cada turno só paga os tokens novos. Muda de modelo ou de máquina e essa memória desaparece, o contexto inteiro é processado outra vez. O @CompleteSkeptic, CEO da @typesafeai, fez as contas numa lista a que chamou a "tirania da KV cache". Ponto 1: "Routing doesn't work." Preços por milhão de tokens, entrada/saída: Opus 5/25, Sonnet 3/15. X é o contexto já na sessão, Y o output, Z o que o agente lê pelo caminho. Ficar no Opus custa 25 Y + 5 Z, porque o X já está em cache. Ir ao Sonnet e voltar custa 3 X + 20 Y + 8 Z: o Sonnet carrega o contexto todo e o Opus ainda tem de ler o que o Sonnet produziu. Com X = 0,65, Y = 0,12 e Z = 0,23 dá 4,15 contra 6,19, cerca de 49% mais caro. As proporções são um palpite, ele próprio diz que as pediu ao ChatGPT. E quanto mais longa a sessão, pior fica o routing. Vi o mesmo na interface da Amália. Tínhamos duas máquinas e o backend alternava entre elas a cada geração, por isso duas sessões seguidas da mesma conversa podiam cair em máquinas diferentes e cada uma reprocessava o prompt inteiro a frio. A correção foi fazer hash do id da conversa e mapeá-lo para uma máquina fixa. Cada conversa passou a aproveitar a cache que já tinha construído. No Hermes, com os meus agentes pessoais, escolho o modelo por tarefa e cada tarefa agendada arranca com um contexto novo e curto, cerca de 6 mil tokens, por isso nunca há cache para perder. Na prática: escolhe o modelo antes de começar a tarefa e mantém a conversa no mesmo modelo e na mesma máquina até ao fim. Se quiseres trocar para um modelo mais barato, troca quando uma tarefa acaba e a seguinte começa, que é quando ainda não há contexto acumulado.
1
2
185
As notas do Diogo Almeida sobre agentes de código, com as contas completas 👇
sharing some notes on typesafe 🤝 coding agents: docs.google.com/document/d/1… we likely will never have time (ever again) to play ourselves, but hope the that the community goes WILD (and makes me look like a naive idiot)
79
my wife after i roll out a manchester triage system simulator with jev for her training
71
Pus o Jev a jogar Batalha Naval e o resultado mais útil é aquele em que ele perdeu contra 50 linhas de código. O Jev, da @typesafeai, responde com probabilidades: envias um estado e perguntas tipadas e ele devolve a probabilidade de cada opção que definiste. Cinco estratégias nos mesmos 60 tabuleiros. Média de tiros para afundar a frota, menos é mais: → Jev híbrido 46,0 → Solver de densidade 48,3 → Heurística hunt/target 51,9 → Jev puro 85,5 → Aleatório 95,3 No Jev puro o modelo recebe o tabuleiro em bruto e as cerca de 90 casas por tentar como opções. Ficou à frente do aleatório, por isso lê o tabuleiro. A heurística ganhou-lhe por 33,6 tiros e em 59 dos 60 tabuleiros. A documentação da TypeSafe explica: as opções têm todas o mesmo rótulo, por isso a lista não traz sinal, e as probabilidades vêm arredondadas a duas casas decimais, o que com 90 opções deixa a maior parte da distribuição a zero. No híbrido o código ordena as casas com um solver de densidade (em quantas posições cada navio ainda cabe em cada casa), descreve as 16 melhores por palavras e o modelo escolhe entre essas. Bate a heurística com p = 0,00015. O solver sozinho faz 48,3, uma diferença dentro do ruído: p = 0,13 e 31 dos 60 tabuleiros para o modelo. O Jev iguala o baseline de código. A configuração em que o modelo fica bem é a mesma em que o código fez a parte difícil. Para quem quiser usar o Jev: lista curta de opções descritas por palavras, contas no código e o baseline de código na mesma tabela. Está tudo open source, com 228 testes, e npm run bench reproduz a tabela. O artigo, os resultados e o repositório estão na resposta.
1
2
183
Lisbon AI Week 🙌

ALT Happy Jeff Goldblum GIF by Apartments.com

1
93
Encontros Mágicos sempre incrível 🪄
72
I benchmarked @typesafeai's Jev at Battleship. The most useful result is the one where it lost to 50 lines of code. Five strategies on the same 60 seeded fleet layouts, 60 games per strategy. 6,000 model calls, $0.77 total. Mean shots to sink the fleet, lower is better: → Jev hybrid 46.0 → Density solver 48.3 → Hunt/Target heuristic 51.9 → Jev pure 85.5 → Random 95.3 Both Jev rows are the same model, jev-1.13.0. The difference between them is what the code hands it. Jev takes a state and typed questions (I used Choice) and returns probabilities over options you define. Jev pure gets the raw board plus all ~90 untried cells as options. Code supplies only the rules. It scored 85.5 shots against 95.3 for random, p = 1.3e-10, so it does read the board. A ~50-line hunt/target heuristic beat it by 33.6 shots and won 59 of 60 boards. ~297K tokens per game. TypeSafe's own docs explain why. Every option carries the same label, "untried cell", so the option list has no signal. Probabilities come back rounded to two decimals, so across ~90 options most of the distribution rounds to zero. Jev hybrid: code ranks the top 16 cells by placement density and describes each one in plain words. The model picks among those. 46.0 shots at ~54K tokens per game. It beats hunt/target with p = 0.00015. The density solver that builds that shortlist scores 48.3 on its own. Against it, 46.0 vs 48.3 is within noise: the confidence interval spans zero, p = 0.13, and the model won 31 of 60 boards. The honest claim is that Jev matches the code baseline. If you're building with Jev: give it a short list of options, each one described in words. Keep the arithmetic in code. Put the code baseline on the same table as the model. From the writeup: "The configuration where the model looks good is the same configuration where the code did the hard part." One methodology finding: the fleet layout family decides the ranking. Density beats hunt/target by 8.3 shots on uniform-random layouts and loses to it by 3.1 on adversarial ones, because density assumes a uniform placement prior. Benchmarking only on random layouts tests the baseline on the distribution it expects.
1
206
This is the follow-through on my "time to try jev" post. Everything is open source: engine, five strategies, benchmark runners. 228 tests, npm run bench reproduces the table. Results: ickas.dev/work/battleship-vs… Repo: github.com/ickas/battleship-… Writeup: ickas.dev/writing/benchmarki… @CompleteSkeptic if there's a better state or option design for this, send it over and I'll rerun it.
68
wake up, baby. it’s time to build.
Jev by @typesafeai is free on Vercel AI Gateway until Sept 25. Build with the fastest adopted model on the Gateway at no cost. vercel.com/ai-gateway/models…
122
We're adding support for AGENTS.md to Claude Code. Starting today in version 2.1.277, if there is no CLAUDE.md in a folder, Claude will check for and use AGENTS.md. You can toggle this behavior in /config.
121
SOLD OUT
Replying to @ickasdev
As últimas vagas estão aqui 👇 luma.com/34od3n97
89