On a mission to apply Python in machine learning and help everyone learn Chinese language.

Dublin City, Ireland
Hello everyone ! I’m on a mission to learn machine learning using Python and also help people to learn Chinese language #machinelearning #Python #Chinese #汉语
1
3
Danny .I retweeted
It is very cool to see llama.cpp on the big stage in todays Windows event! The software and hardware stacks are finally coming together. Our community has put a lot of hard work in the past years and it shows. Looking forward to more users embracing local AI.
24
34
467
12,199
Danny .I retweeted
Google releases EmbeddingGemma 2, a new open model that runs locally on 0.5GB RAM. The 740M parameter Apache 2.0 model combines a 270M text model with vision (170M) + audio (300M). Run & train the model via Unsloth. GGUF: huggingface.co/unsloth/embed… Guide: unsloth.ai/docs/models/embed…
Meet EmbeddingGemma 2, our first natively multimodal open model for on-device embeddings. It expands beyond text to unify code, images, audio, and video in a shared space. 🧵
73
399
3,545
177,831
Danny .I retweeted
Perfect for Multimodal RAG. This is an awesome release. Will add support in localgpt
Introducing EmbeddingGemma 2, a new open multimodal model that sets the standard for on-device efficiency. - our first open, natively multimodal embedding model - handles text, code, image, video, and audio tasks within a lightweight, modular 740M parameter form factor - ideal for offline, privacy-first RAG when paired with Gemma 4 - outperforms some specialist models more than twice its size Weights available now on Hugging Face.
1
5
361
Danny .I retweeted
Introducing EmbeddingGemma 2! 🚀 Our lightweight, multimodal embedding model maps text, code, images, video, and audio into a single, unified embedding space. Optimized for on-device use cases, it features: - 740M parameter form factor with modular encoders - Flexible dimension sizes (768dim-128dim) via Matryoshka Representation Learning (MRL) - 8K context window (4x larger than text-only EmbeddingGemma) - A commercially permissive Apache 2.0 license
97
359
3,420
197,545
Instead of summarising context let the Model edit its memory to preserve context for long runs. This is the idea behind @Meta’s new paper. The LLMs manage their own context natively arxiv.org/html/2609.37725v1 #CLMs
5
Danny .I retweeted
Agent harness tier list.
178
53
1,452
181,785
Danny .I retweeted
We turned Claude Code, Codex, Hermes, Pi, @opencode and other coding harnesses into RL environments. No changes to the harnesses, no changes to the training code. Any open model, any task set, fully open source my friends! Same model, same weights: 62% under Mini-SWE-Agent, 33% under Claude Code. But training inside a real harness normally means reimplementing it as an environment, so most models get trained in a scaffold nobody actually ships. The fix is a proxy, not a rewrite. The harness thinks it's talking to a model API. It's actually talking to a capture proxy that speaks the 4 formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini), forwards to @vllm_project, records the exact token IDs and logprobs vLLM sampled, and hands TRL sequences it can train on. The harness becomes the environment. 10 harnesses run through it today, none modified. And because you control the reward, you can shape behavior the harness never asked for. We added a small bonus for solving a task in fewer tool calls: on tasks it already solved, the model now uses 31% fewer calls, in every harness, and about half under Codex. Tested on LFM2.5-2.6B from @liquidai: → Train in one harness: better mostly in that harness (OpenCode 34% → 58%). → Train in 4 at once: better in all 4 (42% → 54%). → SFT on 3,189 rollouts from Qwen3.8-27B instead: plateaus at 47.5%, below both RL runs. Everything is open and reproducible: the capture proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, the training code and all 7 trained models. Bigger models and bigger runs next. Full guide: huggingface.co/spaces/FineEn…
226
325
2,681
157,894
Danny .I retweeted
Nothing to change here.
147
73
1,706
267,936
Danny .I retweeted
In my opinion, the best model on the planet at the moment is GLM 5.3 Flash. It's super fast, super smart for the size, and is within reach to run locally for many. Love working with GLM 5.3 Flash.
People seriously underestimate how powerful not only GLM 5.3 but GLM 5.3 Flash Crazy intelligence on such a small footprint (yes, small, as in not an entire data center is needed)
107
57
1,694
101,059
Danny .I retweeted
The same model, with the same weights, scores 62% in one agent harness and 33% in another. @adithya_s_k and the @huggingface team just released the ultimate guide to multi-harness RL, and it's one of the most practical RL write-ups this year, and everything open! The trick is simple. Don't touch the harness. Point it at a proxy instead of the model. The proxy speaks all four API formats coding agents use (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages, Gemini). It records the exact token ids and logprobs vLLM sampled, and you train on that. You don't change a single line of Claude Code, Codex or OpenCode. Results: 🔹 Trained across 4 harnesses at once, LFM2.5-2.6B by @liquidai went from 42% to 54% 🔹 31% fewer tool calls, thanks to a small bonus for solving tasks in fewer steps 🔹 Training in OpenCode alone took OpenCode from 34% to 58%, but the multi-harness model improved everywhere They also tried the shortcut everyone reaches for: fine-tune on 3,189 successful rollouts from Qwen3.8-27B. Imitation plateaued at 47.5%, below both RL runs. Copying a bigger model doesn't get you there. Practice does. The best part is that everything is open: the capture proxy in OpenEnv, the trainer in TRL, the tasks, the SFT data, the training code and all seven trained models. Agents will run in dozens of harnesses. Now open models can be trained for each of them, by anyone. Read it here 👇 huggingface.co/spaces/FineEn…
172
162
1,218
119,290
Danny .I retweeted
name a single database better than PostgreSQL I'll wait
168
32
764
41,459
Danny .I retweeted
Anthropic just published its 𝗳𝘂𝗹𝗹 𝗴𝗲𝘁𝘁𝗶𝗻𝗴-𝘀𝘁𝗮𝗿𝘁𝗲𝗱 𝗴𝘂𝗶𝗱𝗲 for Claude Code mods. You can ask Opus 5.5 to use it to go through your past sessions and find the mods you most need to build for yourself. A mod is a small feature you add to Claude Code. The guide's three mods take the things you'd normally stop to ask it and 𝗽𝘂𝘁 𝘁𝗵𝗲𝗺 𝗿𝗶𝗴𝗵𝘁 𝗼𝗻 𝘀𝗰𝗿𝗲𝗲𝗻: how full the context is, what this command is about to delete, and what it just changed. You just describe what you want to see, and Opus 5.5 writes the code. So don't rush to copy the examples. Start with the things you keep stopping to ask about. Send this prompt to Opus 5.5 👇 "Read this guide: claude.dev/blog/getting-star… Then use it to go through my last 20 sessions (this project's .jsonl logs under ~/.claude/projects/) and see how I use you day to day. Find three kinds of things worth turning into a mod: 1. Things I keep asking you, or asking you to do. Each mod in the guide answers one question, like how full the context is. Pick out the questions and requests I repeat across sessions, with counts and my exact words, then decide whether each fits a readout above the prompt, a pane beside the conversation, or a /command. 2. Commands that make me nervous: ones I stopped partway, rejected, or had you undo afterwards. 3. How I check your changes: asking you to list the files you changed, paste the diff, or explain a change. For each idea, give it a name, quote my words, and write a few 'What it should show' lines in the style of the guide's token-weather prompt. If a status line or a setting can already do it, say so, and don't force it into a mod. Rank the ideas by how much they'd help me, and tell me how to keep a mod around once it's built. Show me the ideas first. Don't build anything until I pick."
You can now mod Claude Code: - Change how it behaves - Customize the UI - Swap in your own features Write one with a few lines of TypeScript, or have Claude build it for you. Mods ship inside plugins, so you install them with /plugin in the CLI or desktop app. A few examples:
17
61
1,122
179,439
Danny .I retweeted
Decision models now run on device in llama.cpp. Free, fast, private! llama serve -hf ggml-org/Kev-4B-GGUF
82
121
1,376
62,282
Danny .I retweeted
Decision models in llama.cpp are now available The `/v1/systemone` endpoint is available in the latest llama builds. Use it to do Jev-style inference locally, efficiently and privately. Multiple open models are supported with more to come. huggingface.co/blog/ggml-org…
114
365
2,269
113,804
Danny .I retweeted
.@Cloudflare's decision models are available on Ollama! Decision models allow you to classify an image, label a bug report or route a support ticket to the right team. Clef (27B): ollama pull clef Clef Flash (9B): ollama pull clef-flash
73
139
1,226
75,300
Danny .I retweeted
I don't want an AI that can spend 8 minutes helping me book a flight I want 8 agents spending 8 hours doing work while I sleep like in Matrix @joinmatrixos
2
9
272
Danny .I retweeted
people are building over here in tech europe hackathon 🔥🔥🔥
Enough flight booking discourse We’re back to making agents useful 👀 Matrix maxxing today at Agentic AI Hack in Stockholm with @techeurope_ and @GoogleDeepMind
1
1
18
585