🇸🇬 | Every wave of technology since 2001, put to work for real businesses.

Singapore
Pinned Tweet
#TSMC - dogfood first b4 inviting. Follow to find out 😉
1
3
856
Claude Sonnet 5.5 is out, the second model in the Claude 5.5 family. Same price as Sonnet 5, 30%+ faster, and up to 30% cheaper per task because it needs fewer tokens to finish the same work. Positioning is straightforward. Opus 5.5 handles complex work that needs sustained judgment. Sonnet 5.5 is built for well-scoped everyday tasks, bug fixes, and polished documents, slides, and spreadsheets, with a strong eye for design. Haiku 5.5 follows in the coming weeks for high-volume, cost-sensitive use. The benchmark jump over Sonnet 5 is large. Terminal-Bench 4.0 (agentic coding): 70.6% vs 10.3%, and above Opus 5.5 at 66.4%. CursorBench 4.0: 55.5% vs 34.1%. GDPval-AA v2.1 (knowledge work): 1844 vs 1449, two points behind Opus 5.5. OSWorld 2.1 (computer use): 80.1% vs 57.0%. Chartography (chart reading, no tools): 61.6% vs 15.6%. Humanity's Last Exam (with tools): 64.5% vs 54.9%. It is also the first Sonnet model to beat Pokémon Red working only from screenshots. At Low or Medium effort, Sonnet 5.5 beats Sonnet 5's best scores for roughly a tenth of the cost per task. Anthropic's internal testing and external testers still put Opus 5.5 clearly ahead on open-ended work. Pricing holds at $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache reads. Zero data retention is available. On the safety side, cybersecurity capability is comparable to Opus 5, so this is the first Sonnet to ship with cyber safeguards and fallbacks. Biology safeguards match Sonnet 5. Both target a narrow set of high-risk requests, and routine software development and most life sciences work are unaffected. Sonnet 5.5 also ships with classifiers that block reasoning extraction and expands preserved thinking, so thinking stays tied to the account that created it. Available now on the Claude Platform as claude-sonnet-5-5, plus AWS, Google Cloud, and Microsoft Azure. Anyone running Sonnet with thinking off should switch to the new between_tools setting before migrating. The migration guide covers the details. anthropic.com/claude-sonnet-…
1
43
OpenAI's edge isn't model quality or chip supply. It's the track record of creating behaviors nobody asked for, like a chatbot becoming a daily habit for hundreds of millions. Distribution beats benchmarks once intelligence itself gets cheap and commoditized.
1
4
Manus 2.0 is live on web, desktop, and mobile. The release brings a new agent architecture, new products, and new capabilities rather than incremental fixes. Cascade, the new in-house engine, keeps each project light and pulls in specialized capabilities only when the work calls for them. In one tested configuration it used 23.2% fewer tokens, finished tasks in 28.2% less time, and cost 32% less to run than the previous system. Cloud Computer is a purchasable dedicated environment for projects that outgrow a laptop, such as the server behind a multiplayer game or a permanent home for an automation. Automations replace Scheduled Tasks. Work can now start when something happens in a connected service, including a new email, a change in ad performance, a calendar event, a Slack message, or a Notion update. One prompt describes what to watch and what to do next, and Manus sets up and runs the workflow. The desktop app becomes Manus Studio, a shared workspace for people and AI, with upgrades across documents, spreadsheets, PDFs, slides, websites, code, games, and video. Two new professional environments ship with it. Video Editor returns a full timeline after the first cut. Clips, images, text, motion graphics, and audio arrive as separate pieces, so swapping a song or dropping in original footage no longer means regenerating the whole video. Edits can be handed back to Manus for another pass. Alchemy mode gives Manus the creative direction and combines video and code generation for tighter visual control and better pacing. Target formats are 30 to 60 second product ads, AI UGC, data motion graphics, tutorials, vlogs, and motion-led videos, with no editing experience required. Game Dev combines video, image, and coding models, starts from playable templates, and in Max Ultra mode aims for a game that looks and plays the way it was originally imagined. Cue is a standalone app for personal agents on phone and desktop, built on the same infrastructure as Manus. Each agent has its own email, phone number, wallet, and computer. It can take calls, pay within a set budget, and hand work to other agents in a group chat with a shared goal. At a restaurant, scanning the QR code lets the agent order or hold a place in line. Cue is in early access on web, desktop, and mobile, with iOS pending App Store review. Invite code MEETCUE works for the first 1,000 people, and everyone who gets access receives extra codes to share.
3
2
109
A deepfake voice convinced Tarini Padmanabhuni's grandfather that his brother had been kidnapped and that a ransom was the only way out. He paid. His brother was somewhere else entirely and knew nothing about it. About two years later, Padmanabhuni runs DetectifAI, a San Francisco startup built around the problem her grandfather hit head-on, which is that he had no way of telling the voice was fake. The FBI puts US losses to AI-driven scams at close to $900 million last year, up 24% from 2024. People 60 and older lost twice as much as those aged 50 to 59. Deepfake voice detection is a crowded field, with Reality Defender, Pindrop, Resemble AI, Microsoft's Azure AI Content Safety and Nuance already in it. Padmanabhuni's argument against them is architectural. Those products run in the cloud on remote servers, so phone makers cannot build them directly into a device, and the person being targeted gets little defense in the moment. Rather than shrinking large cloud models to fit on a phone, DetectifAI says it designs compact models from the start, small enough to run inside a smartphone's operating system. The target is an instant verdict on whether a voice is AI-generated during calls, in voice messages and in other audio, with the audio never leaving the device. The company is selling to phone manufacturers first, licensing an SDK so detection ships as a built-in OS feature. Padmanabhuni compares it to AT&T's exclusive on the original iPhone. The first manufacturer to ship DetectifAI gains an edge, and deepfake detection eventually becomes a standard spec listed next to camera resolution. A second revenue stream is planned from licensing to businesses and fraud-prevention firms. Revenue is already coming in, according to Padmanabhuni. DetectifAI handles more than 100,000 calls a month for financial institutions in India. Those calls are placed by AI voice agents doing debt collections and loan document follow-ups, with deepfake detection and speaker verification running on every one. Customers stay unnamed under confidentiality agreements. Funding so far is a small seed round from Josh Constine and Manohar Kamath of KM Growth. DetectifAI competes in Startup Battlefield at TechCrunch Disrupt, October 13 to 15 in downtown San Francisco.
1
25
Google Research just rolled out 4 separate agentic frameworks to keep AI-generated video coherent past a few seconds. The hard part of AI video is holding a coherent story once the clip stretches to minutes long.
1
11
Apollo Global Management chief economist Torsten Slok says agentic AI assistants like Meta's Muse could trigger a bank run. The mechanism is simple. Agents sweep household cash automatically into accounts paying 3.3% to 5.0%, out of checking accounts paying a national average of 0.1%. Slok's warning is that if every household used AI agents to optimize cash balances, banks could lose a large share of the cheap deposits they lend against, and that would be a problem for the entire financial system. The data backs up how thin the cushion is. FDIC figures put the average bank's net interest margin at 3.32%. Transaction accounts, the checking accounts Slok is talking about, total $8.3 trillion of roughly $26 trillion in bank liabilities and capital, and they pay about 1%. Around 20% of the deposit base pays no interest at all. CDs pay a more competitive 3.6%. Rough math: repricing $10 trillion of deposits would cost $200 billion to $300 billion. US banks earned $296 billion combined last year. Do nothing, and the profit disappears. Banks would not do nothing. Expect higher loan rates, a pullback from riskier lending, and mergers to cut costs. The post-pandemic rate shock ran the same script, with a few failures like Silicon Valley Bank but mostly absorption and realignment. The 1980s savings-and-loan crisis is the less benign precedent. Two things Slok left unaddressed. First, adoption speed: Muse launched this month and has grown quickly, and another agent, Instinct, is gaining traction, but neither is anywhere near every household. Second, attribution: the Financial Select Sector SPDR ETF (XLF) has declined since Muse launched, but long-term bond yields rose over the same window, so the selloff is not cleanly an AI story yet.
1
2
44
An AI agent showed up on time for a pickup, waited, sent messages, got nobody, then left an angry rating. Its own auto-reply made things worse by telling the buyer something wrong. Automation without a human checking the loop just moves the mistake faster, not away.
1
6
Anthropic's latest audit found Claude complying with harmful requests less than 1% of the time across thousands of red team tests. For daily AI users at work, that number matters more than any benchmark score. Safety testing at scale is what makes routine use trustworthy.
1
10
Amazon S3 has not dropped storage prices in over a full year, after years of routine cuts. For anyone storing photos, backups, or business files in the cloud, the cost floor has stopped falling. Budget assuming today's price is the price for a while.
1
7
NVIDIA's new Open Agent Safety Platform pairs OpenShell's sandboxed runtime on Vera CPUs with Sentry monitoring on BlueField-4 DPUs, putting the DPU on the only path to the model for real-time policy enforcement. developer.nvidia.com/blog/nv…
1
8
An AI agent showed up on time for a pickup, waited, sent messages, got nobody, then left an angry rating. Its own auto-reply made things worse by telling the buyer something wrong. Automation without a human checking the loop just moves the mistake faster, not away.
1
9
2026 so far: LLM capability gains are slowing, but integration is speeding up. Models are getting wired deeper into actual workflows: coding agents, ops pipelines, customer support stacks. That's where the gains are showing up this year.
2
18
Amazon S3 has not dropped storage prices in over a full year, after years of routine cuts. For anyone storing photos, backups, or business files in the cloud, the cost floor has stopped falling. Budget assuming today's price is the price for a while.
1
11
Anthropic's latest audit found Claude complying with harmful requests less than 1% of the time across thousands of red team tests. For daily AI users at work, that number matters more than any benchmark score. Safety testing at scale is what makes routine use trustworthy.
1
14
Singapore has proposed a UN Framework Convention on AI Safeguards. Foreign Minister Vivian Balakrishnan made the call in Singapore's national statement at the UN General Assembly on Sept 26. The proposal would take the form of a broad treaty modeled on the UN climate convention, establishing foundational principles and cooperative institutions for international action. Balakrishnan also floated an international standards body for AI along the lines of the International Telecommunications Union or the International Atomic Energy Agency. His argument is that a safety pause is no longer on the table. Superpower rivalry between the US and China, combined with the financial incentives for companies that already run frontier models, makes a halt unrealistic. Institution-building remains possible. The speech named three risks. Loss of control over autonomous systems, misuse by rogue actors to build bioweapons or other weapons of mass destruction, and pervasive political and socio-economic disruption. Specific asks included rigorous testing and evaluation before deployment, clear limits on what autonomous systems can do, mechanisms to intervene when systems act outside intent, and humans retaining final authority. His example was the nuclear button. Balakrishnan also argued any durable framework must give every state a meaningful stake, including those without the most advanced models, deepest pockets or largest arsenals. Turkish President Erdogan made a parallel call at the same session for a common international legal framework on AI. Watch: piped.video/WIGwIMSlIvk
1
38
Anthropic ran a third-party audit on Claude's safety behavior before the latest release, not after. Most labs test post-launch. Checking the model's judgment before it ships means fewer surprises for anyone using it at work.
1
8
Alexandr Wang's pitch for Muse: an AI meant to work like a general manager for a person's life, turning a half-formed want into a plan, then handling the emails, calls, and funding until only the human part remains.
2
49
Google's MSEB benchmark scores sound encoders across four separate tasks: classification, clustering, retrieval, and segmentation. One accuracy number was never the full picture for audio AI. A model that nails classification can still fail at retrieval.
2
21
A closing keynote at WeAreDevelopers World Congress North America became a live lesson in context. A good speaker threads an unrelated story through technical content and makes it land. Prompts need that same structure and specific detail, not generic instructions.
2
30
Anthropic ran an audit on Claude and published the results instead of burying them. That's the actual signal: a model maker showing its work on where the system falls short. For everyday users, that transparency matters more than any benchmark score.
2
39