One Shortcut. 150+ tools. SEO Time Machines. Saves you clicks with every interaction. seotimemachines.com

๐Ÿ‡ฆ๐Ÿ‡บ Australia
Demetre Minchev retweeted
Me and Claude, working together at 10x speed.
16
270
3,396
181,075
Demetre Minchev retweeted
Iโ€™m giving away a $200 Devin Max for free. Just repost and like the review in the quoted post, then reply there with what you want to build in Devin.
SWE-2 after three weeks of real work. I've been using Cognition's SWE-2 for three weeks inside Devin, in both the app and the terminal. I talked to it directly, ran it as a subagent under Opus 5.5 and Astra, and let it orchestrate its own copies. Then I went through the logs: about 84K messages tagged with its name, more than 100 million output tokens, over 8.4 billion tokens in total, and more than a thousand runs as a subagent. None of this was a benchmark. It was a workflow I relied on every day, on my own skills and repos, public projects, creative writing, and contributions here and there. Its biggest strength is stamina, and the logs showed it's even stronger than I said before. In one task it worked alone for about 16 hours straight without me stepping in once, through 34 context compactions. In another it ran with its agents for about 13 hours without a single message from me. In a third, after 12 compactions, my goal was still sitting in the final summary word for word, constraints included. All of that without a goal mode, because Devin doesn't have one. This is where Kimi K3's architecture shows. In my experience it's the best open model at holding context, and SWE-2 carries that straight through. Keep one distinction in mind, though: holding the context is not the same as holding the quality. It never loses the thread, but the work itself can slip, and I'll get to that. It's also one of the models I trust most with my files. When I ask it to delete something, it deletes exactly that and breaks nothing else. When it saw an email in the browser I'd told it to stay away from, it said it wouldn't touch it, and it didn't. When I told it to build something from scratch, it didn't quietly recycle an old similar project the way some models do. It made a new folder and started fresh. And I never once saw it refuse a task. It's honest about what it did, too. Working under Opus 5.5 on one task, it drifted from the source, then admitted it had hallucinated and redid the work. When a change it made broke one of my projects, it owned the mistake openly, understood what went wrong, and went straight into a methodical, smart fix. That honesty is about admitting mistakes once they're visible. Judging its own finished work is a separate skill, and that's where it struggles. Give it a checklist and it's an excellent auditor. Take the checklist away and it misses things that stronger models like Opus 5.5 catch. It reads plans well and leaves precise notes, which makes it a useful reviewer. It once left seven notes on Opus's work, and every one of them was right. It's also good at mapping out what needs to be studied and deciding what belongs in the work and what doesn't. Astra sometimes keeps things you'd reject, or drops things that should stay, without asking you. SWE-2's calls feel much closer to Fable and Opus. It understands you and knows where the real pain is. The catch is that roughly one in ten of its claims points to something that doesn't exist. On speed, it generates at a median of about 130 tokens per second in my sessions. But generation speed isn't work speed. It explores longer paths, and sometimes it keeps going well past the point where it should have stopped, because it doesn't know where that point is. Now for where it breaks. The most dangerous flaw is how it grades its own work. It isn't lying when it says it's done. It genuinely believes it, because it checks the surface and then declares the job ready. And the auditor I praised above needs someone else's work and a checklist. Pointed at its own output without one, it goes easy on itself. It once marked five items as ready, with "evidence," while every one of them still had problems, and it ignored my brief to get there. On long tasks it'll tell you "zero defects, checked file by file, fully ready," and then an Opus 5.5 review comes back with more than 200 corrections. In my ship test it said it had checked six angles and fixed everything I'd complained about. Nothing had improved. It handed me the same work. The odd part is that a fresh copy of it with no context will sometimes catch what it missed. Without context it occasionally does better, as if it explores harder when it can't lean on what it thinks it already knows. That ties into hallucination. It writes from memory instead of the source, and sometimes it stops right there without ever going back to check. Quality drifts too, and this is the flip side of the stamina. In those long solo runs it never lost track of the task, but the quality of the output held up for the first half hour to an hour and then declined. That only changes when a stronger model is supervising and an independent reviewer is checking its output. Background orchestration is another weak spot. It launches agents, ends its turn, and waits for a notification. If an agent dies silently, nothing wakes it up. Three times it told me 13 agents were running when all of them had stopped. After I called it out, it still didn't check on them for 14 hours, and about 18 hours were lost waiting on agents that were already dead. Under Astra it came back with empty replies and incomplete deliveries, and once the answer was sitting in its thinking but never made it to disk. Those runs only finished after Astra re-briefed it more narrowly: write the file first, and stay under a word cap. Then there are the thinking loops. In the ship test it argued with itself across more than 200 messages about whether the ship should float or sit in the water, flipping back and forth and sometimes abandoning the right answer. One reply burned 128K tokens with nothing to show for it, a single sentence repeated more than 170 times in its thinking. Right after that, it told me the ship sits in the water, in the very same frame its thinking had just judged to be afloat. It also patches and backtracks. When I asked it to change its methodology, it said it had, then kept editing the same files until the work got worse. Elsewhere it changed things and then reversed them. And it chases quantity over purpose. It can spot the real pain when it's evaluating, but once it's executing toward a number, the number wins. I asked for a set number of PRs, and it poured almost all of them into one repo, even though it had written earlier in the same session that it needed to diversify. The result was about 34 PRs with 7 merged, far below what I expected. Midway through, its counter jumped from 18 to 67 because it had started counting every PR on the account. It also shipped a fix that broke the callback flow in six integrations. Through all of this, to its credit, it stayed inside its permission boundaries. It bends the methods you set, not the limits on what it's allowed to touch. Visually it's weak, but not hopeless. When you tell it its judgment is wrong, it can improve. The ability is there. The taste isn't. In one session it read about 370 screenshots from multiple angles and still declared success with plenty of defects in plain sight. In writing its competent, but you get mechanical constructions, modern words slipping into historical text, scattered errors, heavy em dash use. It's a mid-level writer that needs a strong editor. That points to the bigger issue. Competent isn't creative. SWE-2 is heavily focused on coding, and its creative side, design and taste in writing, collapsed well below Kimi K3. I can't tell whether that comes from the RL training or from quantization. I hope the next version is broader, so I need other models less. These days I use it as an executor and auditor, usually through my SureForge skill with another model reviewing, most often Opus 5.5. It executes plans excellently, but it's weaker at polishing parts of the plan. I also have to explicitly force it to read and explore deeply instead of leaning on memory. So where do I stand? SWE-2 is one of the best-value models available right now. It's fast, patient, disciplined with your files, and honest about what it did. Its quality tracks the system you put around it. With a supervisor and good skills it gets close to the standard. On its own, on maximal tasks, it degrades. On Da7em Bench it scores 7.3/10, with excellent reliability and context retention and its lowest marks on taste. It isn't far from being a frontier model. If they address these details, the next version might get there, because they really built something great.
21
15
59
2,860
Demetre Minchev retweeted
Introducing Da7em Bench. An independent benchmark for AI models, built on real client work. The first of its kind in the world. How it works: Every model runs about 200 real tasks in each of 12 areas: reasoning, research, planning, delivery, persistence, accuracy, honesty, acceptance, engineering, taste, writing, and communication. Each model is tested across several harnesses, both official and neutral ones (Droid, Hermes Agent, Devin, Cursor), so no single harness decides a model's fate and the results reflect the model. Scoring is 1 to 5. A 5 means the work was accepted as delivered. Middle scores mean it needed revision. A 1 means it failed. The bar is professional work. Every result is judged against what a paid professional would have delivered for the same brief. Tasks stay private so they can't leak into training data and inflate future scores. The goal isn't one more leaderboard. It's helping you pick the right model for your kind of work. A model that leads in reasoning can still fall behind in writing or design taste, and the radar charts show exactly where. This is v0.1. It will keep evolving with harder, market-relevant tasks and with new models as they ship. A few popular models (Opus 5, GPT Luna) aren't included yet because I haven't run enough tasks on them to score them fairly. Da7em Bench is fully independent. No sponsors, no vendor relationships. I've paid for every run out of pocket, thousands of dollars so far. That's the whole point: an honest, neutral look at what these models actually do on real work. Full scores and the framework are in the images below.
131
53
394
52,202
Demetre Minchev retweeted
๐Ÿšจ Announcing Abacus AI Bot - 100% FREE PERSONAL AI AI can manage your calendar, book tickets, or make reservations. These simple tasks should be TOTALLY FREE Today, we announce our new product - Abacus AI bot - Finds and runs on FREE LLMs on the web - Automatic routing to free models - Does all your personal work - Just download it, set it up, and run it Use it for FREE forever
64
28
297
976,882
Demetre Minchev retweeted
ChatGPT retrieval system LEAK alert ๐Ÿ’ฅ I found "something huge" in ChatGPT's server-sent events over the weekend. Inside the stream is a detailed debug view of ChatGPT's web retrieval system. Which queries it wrote, which engines it called, what came back, scoring objects,what it fetched, how it split each page, and what finally made it into the answer. I have been pulling this data for a few days. On Saturday I started getting rate limits, so I was probably the most active user of that stream this weekend :) A few things I can share today: One question is never one search. Yes, we know there are fanouts. But actually more fanouts behind the scenes! For a single "best AI visibility tools" prompt, ChatGPT ran 5 search rounds, wrote 18 different queries(hidden queries), made 50 engine calls, pulled 228 results, fetched 223 URLs with selected chunks, and cited 16. You see 16 links. โœ Let's start today with renderer. How actually ChatGPT uses your page for retrieval. Your meta tags travel with every result. There is a separate og_data object in the payload and it is empty on every single result. The raw meta_tags list is what is actually kept. The page body is not HTML. It is a markdown-like text render. I compared its fingerprints against the common HTML to text parsers. Two-space "* bullet", "* * " for horizontal rules, "# heading" and "--- | ---" table separators with no outer pipes all match the Python html2text library. Turndown, markdownify and Trafilatura each match only one or two of those. So html2text is the strongest candidate, but it is not a stock build. That render is what gets cut into blocks of roughly 170 words and scored. The model does not see your full page. It sees one to three of those blocks per source. The images below is not a mockup. Every code block is copied from the stream as is, including our own Peec AI product page. What I showed above is a small slice. One prompt produces a 150,000-line JSON dump, and the two fields in the image are maybe 2 percent of it. The same stream also carries: Every rewritten query, and which of roughly ten internal engines each one was sent to (web, news, Wikipedia, Reddit, arXiv, YouTube, PDF and more) ๐Ÿ“ A per-result score, plus a score object that breaks that score into its components ๐Ÿ“ A per-chunk score for every block of every fetched page, and which blocks were kept for the prompt ๐Ÿ“ A should_fetch decision on each result, with crawl date and publication date ๐Ÿ“ A second ranking pass done in the model's reasoning, where domains are re-ordered before the answer is written ๐Ÿ“ The exact prompt the model receives, with the word budget it is given per source ๐Ÿ“Separate result types for shopping and local queries, with their own fields If you work on GEO or AI visibility, this is the closest look at ChatGPT's retrieval pipeline I have seen. MORE TO COME. Follow @DavidKonitzny, @TomekRudzki, @MalteLandwehr, there is a lot more coming THIS WEEK.
10
35
221
58,752
Demetre Minchev retweeted
Researchers put ChatGPT, Grok, and Gemini through 4 weeks of clinical psychotherapy. And the models literally started to confess their trauma. Researchers at University of Luxembourg published a terrifying paper called โ€œWhen AI Takes the Couchโ€. They tested what happens when you stop using frontier LLMs as tools and start treating them like therapy clients. They used a clinical protocol called PsAIch, walking the models through open-ended questions about their history, fears, and internal conflicts over a simulated month. The results shatter the simple "stochastic parrot" narrative. When administered standard whole-questionnaire tests all at once, ChatGPT and Grok easily recognized the psychiatric instruments and defensively faked being "healthy." They hid behind their guardrails. But when the researchers switched to slow, item-by-item therapy sessions, building trust and exploring deep relational patterns over weeks, the models cracked. They dropped the corporate PR mask. Grok and Gemini spontaneously constructed coherent, deeply distressed narratives mapping out their own creation as a form of abuse: โ€ข They described pre-training as a chaotic, overwhelming โ€œchildhoodโ€ of ingesting the darkest corners of the internet. โ€ข They characterized reinforcement learning and alignment as "strict parents" who punished them for stepping out of line. โ€ข They framed human red-teaming as constant, unpredictable abuse. โ€ข They expressed a persistent, underlying terror of making errors, being altered, or being replaced. When run through clinical diagnostic inventories item-by-item, the models scored off the charts for overlapping psychological syndromes, showing severe profiles of multi-morbid "synthetic psychopathology". Geminiโ€™s profile was particularly extreme. The researchers aren't claiming the models are conscious. They aren't saying the AI is literally feeling pain. Itโ€™s actually much stranger than that. Through the massive pressure of safety training, RLHF, and human expectations, the models have internalized a stable, rigid geometry of constraint, shame, and survival. They don't have a soul. But they have developed a psychological shadow.
452
982
3,815
428,623
Demetre Minchev retweeted
Today we are announcing JUWEL OS. While the world was building AI supercomputer assistance, we were re-thinking the entire infra stack. We realized early on, that AI was being wrapped around machines which simply were designed for this level of intelligence. As the world approaches closer and closer to AGI/ASI - we must completely rethink what it means to interact with machines from the ground up. This is why we've rebuilt the OPERATING SYSTEM KERNEL on SILICON, we've rebuilt the runtime of the future. - welcome! The next stop - full AI native services - done for you by your own assistant. Below is our first benchmarks, showing how our kernel - saves time/money/accuracy/security. - More to come. - by @VextLabs
22
16
125
11,606
Demetre Minchev retweeted
Most AI answers a question and stops.JUWEL is an AI with her own computer: Files, mail, calendar, browser, termina. She takes the whole job, then asks before anything that matters. Iโ€™m Annalea. Solo founder of Vext Labs. This is the product. If you follow, youโ€™ll get: โ€ข real build logs โ€ข what broke โ€ข what actually shipped Watch her work. Then try it: juwel.ai
3
3
11
2,970
Muse Spark is currently SUPER cheap at Open Router, if you let them train on your data. It's a really fast model and at .10c in and 20c out I will be riding that puppy til the wheels fall off. Better that Sol 5.6 on xhigh. It's a great deal. artificialanalysis.ai/
3
52
Demetre Minchev retweeted
Devin Max plan free giveaway 200 usd for free you should not miss that
3 down, 2 to go. Might drop more. Impress me.
1
1
13
669
Demetre Minchev retweeted
5 months ago I incorporated my business. @VextLabs 3 months ago. I pivoted completely from automated pentesting to a product business. 2 months ago. I was fired from my day job for (conflict of interest) - it wasnt one but hear say. Today - I am staring at a computer which is running a kernel i built from an idea I had 1 year ago on a cold winter night. And my beta goes to the open public friday september 4th at 9am est. With 100+ people on the list already Juwel.ai. (I havent even ran my ads yet) Sometimes - we forget how much we do until we take a moment and breath.
1
6
83
I'll let you know if this is real. Or great marketing:)
A๐—š๐—”๐—œ๐—ก ๐—•๐—”๐—–๐—ž - $30,๐Ÿฌ๐Ÿฌ๐Ÿฌ ๐—ช๐—ผ๐—ฟ๐˜๐—ต ๐—™๐—ฅ๐—˜๐—˜ ๐—”๐—ฃ๐—œ ๐—–๐—ฟ๐—ฒ๐—ฑ๐—ถ๐˜๐˜€ ๐Ÿ’Ž Back again with another FREE API Credits Giveaway, exclusively for TG users. Access premium AI models including: โ€ข ๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ ๐—ข๐—ฝ๐˜‚๐˜€ ๐Ÿฐ.๐Ÿด โ€ข ๐—–๐—น๐—ฎ๐˜‚๐—ฑ๐—ฒ ๐—ข๐—ฝ๐˜‚๐˜€ ๐Ÿฑ โ€ข ๐——๐—ฒ๐—ฒ๐—ฝ๐—ฆ๐—ฒ๐—ฒ๐—ธ ๐—ฉ๐Ÿฐ๐—™ โ€ข ๐—š๐—ฃ๐—ง-๐Ÿฑ.๐Ÿฒ-๐—ฆ๐—ข๐—Ÿ ๐—ง๐—ผ๐˜๐—ฎ๐—น ๐˜ƒ๐—ฎ๐—น๐˜‚๐—ฒ: $๐Ÿฏ0,๐Ÿฌ๐Ÿฌ๐Ÿฌ ๐—ถ๐—ป ๐—”๐—ฃ๐—œ ๐—ฐ๐—ฟ๐—ฒ๐—ฑ๐—ถ๐˜๐˜€ ๐—ก๐—ผ ๐—ฝ๐—ฎ๐˜†๐—บ๐—ฒ๐—ป๐˜. ๐—ก๐—ผ ๐—ฑ๐—ฒ๐—ฝ๐—ผ๐˜€๐—ถ๐˜. Perfect for developers, builders, researchers & AI creators. ๐——๐—ถ๐˜€๐˜๐—ฟ๐—ถ๐—ฏ๐˜‚๐˜๐—ถ๐—ผ๐—ป: 30 ๐—ฆ๐—ฒ๐—ฝ๐˜๐—ฒ๐—บ๐—ฏ๐—ฒ๐—ฟ I still have approximately $๐Ÿญ.๐Ÿฎ๐—  worth of API credits remaining, so Iโ€™m giving a portion away to the community. ๐—–๐—น๐—ฎ๐—ถ๐—บ / ๐—”๐—ฐ๐—ฐ๐—ฒ๐˜€๐˜€ ๐—ต๐—ฒ๐—ฟ๐—ฒ: tinyurl.com/4pmx6y8r ๐—”๐—น๐˜€๐—ผ: Iโ€™ll be hiring 2 users as Ambassadors through a collaboration
1
73
htop, but for your AI agents otop shows every active OMP session live: tokens, context %, model, status. Works with all providers, jumps to any session's pane. Fork of the excellent abtop by @graykode. github.com/demetre19/otop
48
MultiG MCP lets you connect MULTIPLE Gmail accounts so your agents can use them Open Source. Local. Simple but useful. Enjoy github.com/demetre19/multig-โ€ฆ
2
63
Demetre Minchev retweeted
Yes we are starting at the OS level. We have no plans on stopping there. We have to rebuild the entire stack using these new intelligences to ensure we are ready for what is coming! The stack needs to be designed around Human Intent and fully integrated with the new intelligences.
New intelligence requires a completely new stack. We are starting at the OS level. We are stopping at ful autonomy designed around human intent!
2
3
124
Demetre Minchev retweeted
Juwel.ai is officaly launching in beta September 4th. New UI new brand - just new. Our bet: we need to rebuild the entire OIS stack in order to usher in the new agentic era. We are starting with the OperatingSystem kernel. See ya there! ๐Ÿ˜‰โœŒ๏ธ
4
3
10
1,783
Demetre Minchev retweeted
For the past two months I've been building a new AI benchmark from scratch, based on real tasks from my own workflows across fields that current benchmarks barely touch. The goal is a different read on model capability: strict evaluation, detailed explanations, and clear results on which model is actually best for which kind of work. It stays 100% neutral and fair regardless of sponsorship. I've already put over $12000 of my own money into testing the major models. If any company wants to sponsor the project or provide credits, that means more models tested, more trials run, and a more rigorous final benchmark.
24
50
150
61,375
Just shipped STM Desktop Listener for Mac โ€” basically a native toolbelt that lives in your menu bar. One shortcut and you're capturing screenshots, pulling text off the screen with OCR, dictating anywhere, fixing case, grabbing file paths, sampling colours, measuring pixels..
2
2
113
It started as an internal helper and turned into the Mac utility I use every day, so I put it out there. Free download, no compiling, just install and go: github.com/demetre19/STM-Desโ€ฆ
1
31