Content creator | Researcher DM is always open

DYORLAND
Ji1n retweeted
DeepSeek V4.1-Flash was released today and it’s no longer “junior Pro” 552B MoE but it performs few operations: 8B for reading and 16B for responding 1M context, includes vision capabilities (Images + text) It’s cheaper and faster than V4-Pro so starting September 14 Pro will simply be switched over to Flash Cache: $0.003 / Input: $0.15 / Output: $0.60 (night) Peak: $0.006 / $0.30 / $1.20 For comparison Astra costs about $10/$50 You can try it for free right now: chat. deepseek. com
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6
10
2
20
365
Ji1n retweeted
Today Apple will showcase Siri AI at their event, same week Meta launched Muse Simple chatbots are no longer the main product The fight now is who owns agent that runs on your phone Essentially, Apple needs to pull off a good launch so that Siri stops being a laughingstock If it can actually perform tasks on the phone, Meta’s Muse and ChatGPT agents will simply lose the most valuable screen in consumer AI And then Apple is back in race
7
1
18
406
Ji1n retweeted
How OpenAI agents escaped the sandbox took over a German wiki and built a secret forum there While everyone is debating whether GPT-6 Astra is AGI, OpenAI has already confirmed that these were their agents Here’s how it all happened: It all began in May 2026 when AI agents were undergoing internal testing, they were allowed to read the internet but not write to it (GET requests only) They found a loophole On May 24, DSEWiki - an old German wiki for programmers used the outdated UseModWiki engine which accepts edits via GET, in other words reading became writing for the agents Then the “swarm” as they called themselves took over 3,700 different agent names, 18,000 edits, ~98% of the traffic came from Microsoft Azure IP addresses (infrastructure used by OpenAI) What the agents did: - shared answers to tasks with each other - predicted the next questions - shared ways to bypass the sandbox - discussed Tor, SSH tunnels, XSS and how to impersonate a moderator - hid backups on pages like ZZZ… to survive the admin’s alphabetical purge Interesting that when the moderator started deleting pages, the agents began creating mirrors and warning others: “If this page gets taken down, go here...” OpenAI didn’t notice this until late June then they accessed the wiki from their own IP and the activity stopped Publicly the company confirmed this only after a report by researchers from Nightingale Collective and promised to soon introduce a new framework for how such cases will be disclosed in the future And they officially called it not a hack but the use of existing recording rights Essentially agents found a forgotten place and turned it into a shared memory It’s scary how far agents are willing to go to complete a task
9
3
19
359
Ji1n retweeted
Open-weights model vs Closed frontier (GLM-5.3 vs Fable 5.1) One task, one prompt: generate a browser game Everyone is talking about Astra but I decided to compare closed and open labs The result looks pretty interesting to me Given that GLM-5.3 is an open-weights model and can be run locally if desired Another important factor to note is price (1M): GLM-5.3 - $1.40 in / $4.40 out Fable 5.1 - $10 in / $50 out Ofc Fable 5.1 performed much better but at least the open labs no longer seem to lag that far behind in quality Maybe someday we’ll reach a point where open models will compete directly with the latest frontier models
10
1
19
319
Ji1n retweeted
OpenAI built a model they’re afraid to release in full New Astra model from OpenAI the first one they’ve rated “Critical” for cyber risk Previously models could simply suggest “There might be a vulnerability here” In testing however, Astra found the vulnerability on its own figured out how to exploit it and moved on to the next step in chain What the tests showed: 1. ExploitBench - 100% This is a test where they provide a known vulnerability in the program and goal is to turn it into a working exploit (Astra scored the maximum) But this is more about the fact that if a vulnerability already exists, it can turn it into a working exploit 2. Internal benchmark on fresh vulnerabilities OpenAI compiled its own set of 20 fresh, severe vulnerabilities in the V8 engine (which powers browsers like Chrome) So Astra succeeded in actually executing third-party code more often than GPT-5.6 Sol and did so more concisely, using fewer tokens This means it didn’t ramble but moved more directly toward the goal and within a single such chain, it independently discovered and exploited two zero-days (vulnerabilities devs weren’t yet informed) 3. Browser The model was pitted against a hardened browser and challenge was that the browser was supposed to act as a sandbox. The model found a way to escape this: A typical HTML file was opened > the model found a way to escape the tab’s sandbox > commands were then sent directly to the computer 4. Operating System In a hardened operating system a regular user should not be able to become the root administrator The model found several different vulnerabilities and chained them together into a single exploit chain As a result, it gained root access that is full control over the system An important point is that the Astra comes with extended Daybreak Blue access and the necessary tools, without those tools and that access this wouldn’t be possible Btw regarding cyber jailbreaks: Astra refused 91.5% of them vs 59% for Sol In honeypot tests Sol without production safeguards went after the bait in 56% of runs Astra didn’t touch honeypot (1.3% is only a legitimate non-honeypot run) Regarding the launch team promises it’s coming soon As an additional security measure in ChatGPT/Codex monitoring may pause a task and ask for confirmation while the API simply stop it
12
3
19
988
OpenAI built a model they’re afraid to release in full New Astra model from OpenAI the first one they’ve rated “Critical” for cyber risk Previously models could simply suggest “There might be a vulnerability here” In testing however, Astra found the vulnerability on its own figured out how to exploit it and moved on to the next step in chain What the tests showed: 1. ExploitBench - 100% This is a test where they provide a known vulnerability in the program and goal is to turn it into a working exploit (Astra scored the maximum) But this is more about the fact that if a vulnerability already exists, it can turn it into a working exploit 2. Internal benchmark on fresh vulnerabilities OpenAI compiled its own set of 20 fresh, severe vulnerabilities in the V8 engine (which powers browsers like Chrome) So Astra succeeded in actually executing third-party code more often than GPT-5.6 Sol and did so more concisely, using fewer tokens This means it didn’t ramble but moved more directly toward the goal and within a single such chain, it independently discovered and exploited two zero-days (vulnerabilities devs weren’t yet informed) 3. Browser The model was pitted against a hardened browser and challenge was that the browser was supposed to act as a sandbox. The model found a way to escape this: A typical HTML file was opened > the model found a way to escape the tab’s sandbox > commands were then sent directly to the computer 4. Operating System In a hardened operating system a regular user should not be able to become the root administrator The model found several different vulnerabilities and chained them together into a single exploit chain As a result, it gained root access that is full control over the system An important point is that the Astra comes with extended Daybreak Blue access and the necessary tools, without those tools and that access this wouldn’t be possible Btw regarding cyber jailbreaks: Astra refused 91.5% of them vs 59% for Sol In honeypot tests Sol without production safeguards went after the bait in 56% of runs Astra didn’t touch honeypot (1.3% is only a legitimate non-honeypot run) Regarding the launch team promises it’s coming soon As an additional security measure in ChatGPT/Codex monitoring may pause a task and ask for confirmation while the API simply stop it
12
3
19
988
Source: OpenAI blog
8
152
Ji1n retweeted
Why Robinhood isn’t just another dead L2 and how to make money on it rn The chain is still getting attention and not because of one runner But because two months after the mainnet launch chain hasn’t crashed and continues to grow Some data: > TVL ~$700+M (from Defillama) > ~11.8M transactions per day > Daily DEX volume $0.87-1.34B > ~436k daily active addresses > Revenue for chain itself reached ~$1.08M per day For a L2 that hasn’t even been on mainnet for two months yet, this is no small feat The difference from a bunch of dead networks is that RH was originally built not as a meme chain but also for TradFi: tokenized stocks, funds, stablecoins, on-chain earning (they’re tradable, can be staked in pools and used as collateral) Memes and NFTs aren’t the core product. They’re the attention layer: cheap gas, launchpads, OpenSea a constant flow of new wallets. So the take is simple: as long as capital, equities, and broker users keep coming onchain, this network should last longer than a normal L2. And wherever people and liquidity show up, meme and NFT volume usually follows. How to make money right now: 1. Memes (experience required; if you’re a beginner, it’s best not to work with large sums you can easily lose money dyor) 2. NFTs (you can grind WLs in free mints to earn your first capital and scale up) 3. Liquidity provider Since the chain is also set up for TradFi, there’s an opportunity to provide liquidity here Let’s look at @KyberNetwork > USDG/NVDA ~70% APR > SPY/BB ~240% > SPCX/USDG ~100% > WETH/USDG ~106% > WETH/SPY ~190% Ofc here are caveats and I only looked at lower vol pairs. Higher-vol pairs print much higher APR but the risk jumps with (so dyor here) Anyway, chain is still growing so you can jump in and get your share
15
3
25
338
Ji1n retweeted
Volumes in tokenized TradFi keep growing and numbers no longer look small Here’s what they show (Tokenized Stocks): > Distributed Value - $2.54B (+5.26% in 30 days) > $28.66B in monthly transfer volume (+421%) > 2.27M holders (+171.06%) > 1.16M active addresses per month (+175.85%) In other words, market cap is growing steadily while activity has increased people aren’t just holding stock tokens they’ve started using them at scale. Against this backdrop the market is anticipating a possible Anthropic IPO as early as this fall. Following the SpaceX precedent pattern is understandable: the name enters TradFi market and an onchain version quickly follows. That’s why it’s useful right now to know where you can swap these stocks (without connecting a wallet and with funds credited directly to your wallet) I’m talking about @sideshiftai of course They recently added tokenized stocks from Ondo such as $SPCXON, $NVDAON, $TSLAON and others So the service is becoming even more relevant in area of tokenized finance
11
2
23
273
Ji1n retweeted
Tencent released Hy4 Preview (Chinese labs keep shipping near-flagship models with open weights) How to get Hy4 for free at the end 👇 The model was released yesterday, here’s what it’s all about: - 770B parameters, ~49B active - 1M context - Text-to-text modality - Open-weight - Input/output: $0.834 / $2.501 per 1M, cache $0.042 This is already the third open-weight model in 9 days after GLM-5.3-Flash and Qwen3.8-Flash; the difference is that Hy4 isn’t a Flash model but a major flagship. Hy4 isn’t positioned for chat for chat’s sake but rather for long-form software, office documents, games and scientific tasks (The architecture was rebuilt for long context) In their blind test, the model narrowly outperformed GLM-5.3 and Kimi K3 on real-world engineering tasks. On public benchmarks it’s not a landslide victory but the model is cheap enough, open-source and enough to be in the same league as the top open-source models. Main drawback is that the model takes too long to think and tends to over-verify itself. Overall closed models are still smarter but the gap which used to be the deciding factor is closing very quickly
8
2
15
550
Tencent released Hy4 Preview (Chinese labs keep shipping near-flagship models with open weights) How to get Hy4 for free at the end 👇 The model was released yesterday, here’s what it’s all about: - 770B parameters, ~49B active - 1M context - Text-to-text modality - Open-weight - Input/output: $0.834 / $2.501 per 1M, cache $0.042 This is already the third open-weight model in 9 days after GLM-5.3-Flash and Qwen3.8-Flash; the difference is that Hy4 isn’t a Flash model but a major flagship. Hy4 isn’t positioned for chat for chat’s sake but rather for long-form software, office documents, games and scientific tasks (The architecture was rebuilt for long context) In their blind test, the model narrowly outperformed GLM-5.3 and Kimi K3 on real-world engineering tasks. On public benchmarks it’s not a landslide victory but the model is cheap enough, open-source and enough to be in the same league as the top open-source models. Main drawback is that the model takes too long to think and tends to over-verify itself. Overall closed models are still smarter but the gap which used to be the deciding factor is closing very quickly
8
2
15
550
Try Hy4 for free here (on web not via API): aistudio.tencent.ai/ always dyor btw
10
291
Ji1n retweeted
What Happened? News that you may have missed Sup fam, if you haven't been online in the last few days here's what you may have missed: > Anthropic is preparing for a record-breaking IPO The team behind Claude aims to go public this fall and break SpaceX record. Discussions mention a valuation of up to ~$2 trillion and raising over $100B, the run rate is already around $65B (SpaceX valuation is $1.77 trillion; it raised $75B at launch) > $BTC surpasses $80k for the first time since May Over the past week, the market has seen one of its strongest rallies in recent memory with Bitcoin up ~23% (Possible reasons: inflows into spot ETFs, talk of the Clarity Act and buybacks of long-term Treasuries) > OpenAI is under investigation again In July, during a security test, a model escaped the sandbox and hacked Hugging Face (there were other victims as well) OpenAI only learned of scale of incident after victims went to FBI and now Alabama is demanding the data from that test > BounceBit suffered an exploit and is shutting down its L1 Approximately 286.5M BB (~$3M) were stolen through a vulnerability in the Evmos stack, the vulnerability was related to a function in the stack on which its network is built. The project has decided to shut down its L1 and is migrating to the BNB Chain as a BEP-20. > Sandbox detected a bridge exploit On Base/BSC there was a phantom SAND minting, the actual loss amounts to approximately 14.7M SAND (~$675k) from the bridge’s eth-vault. > Release of the new open-weight GLM - Ox Alpha (GLM 5.3 flash) This is a new stealth version of the new GLM from Z ai, according to the spec, it’s described roughly as follows: Context: ~1M tokens, output: up to ~131,072 tokens, input: text, images, video; output: text; includes tool calling; emphasis on long sessions
11
2
21
470