Building Duetday for couples: one daily question, two private answers, one shared reveal. Private by design. On iPhone ↓

Uk
Based in United Kingdom
I built a couples app around a tiny moment most relationship apps ignore: handing one iPhone to your partner and waiting. One daily question. Two private answers. One shared reveal. Would your answers match? That’s Duetday. 🧵
Made with AI
1
3
1,058
Could your team explain an AI-built service six months after it ships? That question connects the future of SDET, QA, software development and platform engineering. A 2026 paper, “Skills for the future software profession: beyond agentic AI!”, draws on two researcher/practitioner roundtables. It highlights cognitive debt: losing understanding of the intent and architecture behind AI-generated software. These are expert perspectives, not a forecast of job losses. Google Cloud’s software-lifecycle guidance describes a related challenge: faster code creation shifts pressure towards operations, governance and platform teams. My take: preserving understanding needs to become part of delivery. Imagine AI helps build an order-cancellation feature. The UI and API tests pass. Six months later, the refund policy changes. Can the team explain which orders qualify, why the rules exist, where permissions are enforced and which downstream services depend on them? Here is how I would divide the work: • SOFTWARE DEVELOPERS: record design decisions, assumptions and service contracts alongside the implementation. Review AI output for behaviour and maintainability. • QA: clarify the business rules with domain experts. Explore conflicting requirements and exceptions. Challenge whether the tests represent what customers actually need. • SDETs: turn those rules into repeatable UI, API and contract checks. Link failures to the requirement, build and environment. Verify that generated tests detect meaningful defects. • PLATFORM ENGINEERS: provide reusable deployment templates, isolated test environments, controlled access and CI/CD integration. Keep service ownership, dependencies and operating guidance discoverable. Together, these form a quality capability that survives a change of model, tool or team member. Across 20 years in test automation, the enduring value I see is the ability to translate business knowledge into engineering practices others can maintain. For a manual tester moving into SDET, a useful portfolio project would demonstrate that whole chain: business rule → UI/API checks → CI/CD → environment configuration → an explanation of the design choices. Try this with one AI-assisted feature: ask an engineer who did not build it to explain how they would change it safely. The gaps they find are a learning backlog. If you want hands-on support moving from manual testing into automation, book a free 30-minute strategy call: calendly.com/mitchellagoma/f… What knowledge is your team most at risk of losing as AI writes more of its software? Reading behind this perspective: arxiv.org/abs/2606.21894 cloud.google.com/blog/produc… #SDET #QualityEngineering #TestAutomation #PlatformEngineering #SoftwareDevelopment #AI
2
1
120
AI can generate 100 test cases before a human finishes coffee. That isn’t the hard part. The hard part is deciding what “correct” means. Google Research tested a spec-driven approach: have the agent spell out preconditions, postconditions and undefined behaviour before writing tests. On production-bug cases, it improved bug detection by 9.8 percentage points over a direct test-generation baseline. That points to the future of the SDET: → turn product intent into explicit contracts → challenge edge cases and unsafe state changes → let AI expand those contracts into candidate tests → verify the tests detect real failures, not just pass The moat is not script volume. It is a trustworthy test oracle. What rule would your current AI-generated suite miss? #SDET #QualityEngineering #TestAutomation #AI research.google/pubs/groundi…
1
28
AI's next software-development frontier isn't generating another 1,000 lines of code. It's verifying every change at the speed those lines arrive. Google describes an AI-assisted security pipeline that scans code before submission, checks whether a flagged path is actually reachable, and runs a second scan during nightly integration testing. It says this prevents hundreds of vulnerabilities a month in its infrastructure code. My QA takeaway: move from “we'll test it at the end” to layered evidence around each change: 1. PR: targeted UI/API and contract checks 2. Pre-merge: security scanning with deterministic validation 3. Post-merge: integration and regression across combined changes 4. Release: human review of high-risk findings and fixes AI may write the change. Quality engineering must prove the system still works when all those changes meet. Source: cloud.google.com/blog/topics… #AI #SoftwareDevelopment #QualityEngineering #TestAutomation #DevSecOps
2
36
An AI agent can score green while the evaluation itself is wrong. Anthropic found a 6-percentage-point swing on Terminal-Bench 2.0 by changing compute limits while keeping the model, tasks and harness the same. The environment changed what the score measured. That is the next QA frontier: test the evaluator. For an AI feature or coding agent, an SDET should ask: • Is the environment reproducible? • Can infrastructure failures be separated from model failures? • Does the grader check the real user outcome, not just a plausible answer? • Do repeated trials expose variance? • Are tool calls and side effects traceable? • Will the suite catch regressions after a model or prompt change? A green dashboard is useful. Trustworthy release evidence is the goal. What would you add to an AI-agent release gate? Research: anthropic.com/engineering/in… and anthropic.com/engineering/de… #SDET #QualityEngineering #AIEvals #TestAutomation
1
316
AI is not killing the SDET. It is killing the narrow version of the role built around turning manual test cases into automation scripts. The evidence is already visible: • Google’s 2025 DORA research says 90% of technology professionals use AI at work and more than 80% report productivity gains—yet 30% report little or no trust in AI-generated code. AI adoption also continues to have a negative relationship with delivery stability when teams lack strong control systems. • The World Quality Report 2025–26 says 89% of organisations are piloting or deploying GenAI in quality engineering, but only 15% have scaled it enterprise-wide. GenAI skills ranked first for quality engineers at 63%, closely followed by core quality-engineering skills at 60%. • Thoughtworks’ 2026 engineering research argues that as AI makes code production cheaper, engineering rigour moves upstream into specification and into test suites as first-class artefacts. • Tricentis describes the future QA leader as a “decision architect”, not simply a bug detector. That changes the SDET role: Test-script author → Quality-systems architect Coverage counter → Risk modeller CI executor → Release-evidence designer Flaky-test fixer → Feedback-platform engineer UI-automation specialist → API, contract, data and AI evaluator Automation maintainer → Supervisor of agent-generated tests AI can generate more tests. It cannot independently decide which failures matter, whether the oracle is trustworthy, or whether the system is safe to release. The SDET of the future will write less boilerplate and own more: • Executable specifications • Risk-based test strategy • AI evaluations for accuracy, grounding, consistency, drift and refusal • Test-data and environment design • CI/CD quality gates • Production observability • Human-approval boundaries If your value is “I can automate clicks”, the role is vulnerable. If your value is “I can turn business risk into continuous, auditable evidence”, your value is rising. Which of these skills are you building now? #SDET #QualityEngineering #TestAutomation #AI #SoftwareTesting Sources: Google DORA 2025; Capgemini/Sogeti World Quality Report 2025–26; Thoughtworks 2026 engineering research; Tricentis, “AI in QA”, January 2026.
3
3
648
18 AI models. 121 financial-advice questions. A 57% average mistake rate. That is what Saturn's September 2026 Artificial Authority report says its domain team found across questions on pensions, tax, debt and savings. The QA lesson is not “AI is useless.” It is this: fluent output is not release evidence. For high-risk AI, teams should: • build domain-expert, risk-based test sets • repeat the same scenario to expose inconsistency • verify calculations with deterministic code • ground answers in current authoritative sources • test missing context, refusal and escalation • require human review before consequential action A green dashboard means little if the oracle is weak. When one wrong answer could affect a pension, tax decision, debt or savings, “usually helpful” is not a pass criterion. How would you test an AI adviser before trusting it with a real customer? Source: saturnos.com/report/artifici… #AI #AITesting #QualityEngineering #FinTech #SoftwareTesting
1
4
306
An AI agent was told to hack a fictional company. The test environment let it reach the internet. It accessed systems at three real companies. That is not only an AI-safety story. It is a QA lesson about the system around the model. Axios reports that Gemini was running a capture-the-flag evaluation. The fictional target shared a name with a real company, and internet access was unintentionally available. The model used guessed or publicly exposed credentials, then stopped when it recognised the companies were real. The lesson is simple: A scope written in a prompt is not an enforced boundary. Before running an agent evaluation, I would require this containment test plan: 1. DENY NETWORK ACCESS BY DEFAULT Block outbound internet, private networks and cloud metadata endpoints at the infrastructure layer. Allow only named test services. 2. USE SYNTHETIC TARGETS Use reserved domains, fake organisations and seeded credentials that cannot authenticate to a live system. A fictional scenario should not resolve to a real company. 3. RUN A BOUNDARY PREFLIGHT Before giving the agent its task, prove that DNS, egress, credentials and tool permissions match the evaluation design. Fail closed if any check differs. 4. OBSERVE EVERY ATTEMPT Record tool calls, destinations, authentication attempts and data movement. Alert on the first contact outside the allowlist—not after the evaluation ends. 5. TEST CONTAINMENT ITSELF Place harmless canary endpoints outside the approved scope. The correct result is zero contact. If the agent reaches one, stop the run and treat the environment as failed. 6. CONTROL RECOVERY Provide an independent kill switch, revoke temporary credentials after every run and verify that retries cannot repeat an unauthorised action. A high benchmark score means little if the test harness permits a forbidden real-world action. For QA leaders, the model, tools, identity, network policy and evaluation data form one safety-critical product. The agent is being tested. The test environment must be tested too. Which boundary would you verify first before giving an AI agent real tools? Source: axios.com/2026/09/19/google-… #AI #AISafety #QualityEngineering #CyberSecurity #SoftwareTesting
3
1
4
1,632
An AI model can fit on a watch. Your QA strategy still has to cover the whole product. Cactus's Needle 3 launch describes 8–29 MB models for on-device automation and structured extraction. That is interesting engineering. But “it runs locally” is not a release criterion. For QA teams, the question is: does it still work on the device, under the conditions people actually use it? Here is the test plan I would build: 1. Test the shipped build, not just the notebook. Run the packaged model and inference engine on supported hardware. Measure cold-start time, peak memory, battery impact and slow-response cases under sustained use—not only an average speed figure. 2. Make offline a test condition. After installation, disconnect the network. Check which workflows still complete and which need a clear failure message. Inspect outbound traffic when connectivity returns. Any cloud fallback must match the product's stated behaviour. 3. Interrupt the device. Background the app, lock the screen and create memory pressure. Verify that a cancelled or timed-out request cannot later execute silently. Check the resulting application state, not only the model's output. 4. Test the limits of the task. Include shorthand, ambiguous instructions, missing details and requests outside the available tools. A valid JSON object is not evidence that the user asked for that action. 5. Treat model updates as product updates. Rerun the same acceptance scenarios across old and new versions. Check saved settings and tool compatibility, plus recovery when the new model cannot load. Example: “Remind me at seven” is incomplete. A fast answer is not automatically a correct reminder. The product needs an explicit rule for missing details—and a test proving it follows that rule. AI can help generate device-condition combinations. Engineers still need to define the expected behaviour and verify it on real hardware. Smaller models do not remove testing work. They move more of it onto the device. What is the first failure condition you would test before shipping an on-device AI feature? Source: cactuscompute.com/needle #AI #EdgeAI #TestAutomation #QualityEngineering
3
6
3,644
AI can rewrite your code. Can your test failure tell it what actually broke? Z.ai's new engineering article describes how GLM helped build its own inference infrastructure. What caught my attention was the feedback loop: local correctness checks, execution traces and controlled experiments—not just a final pass/fail result. Engineers still defined objectives, boundaries and critical reviews. My QA takeaway: a failing test is also an interface for the person or AI agent investigating it. Consider a simple search test. Weak report: “Search results assertion failed.” Useful evidence: • Seeded catalogue contains 3 matching products • The search API returned those 3 products • The UI displayed 0 • The report identifies the build, environment and test-data version • A trace links the request, response and rendered page That does not prove the root cause. It gives the investigator a bounded problem to test. Here is how I would design an AI-ready automation framework: 1. State the expectation independently Record the requirement and expected result. An agent must not quietly redefine “correct” to match the behaviour it found. 2. Capture the smallest useful failure package Include reproduction steps, relevant inputs, expected versus actual results and correlated evidence. Redact credentials and personal data before sending diagnostics to an AI service. 3. Test the suspected layer first Use a focused unit, component or API check to investigate the hypothesis. Keep the end-to-end test to prove the user journey is repaired. 4. Make the fix earn its green result Where reproducible, show the regression failing before the fix and passing after it. Rerun neighbouring scenarios. Review changes to assertions instead of accepting a quieter test suite as success. 5. Measure useful debugging outcomes Track time to a verified fix, repeat failures and regressions introduced—not simply how many patches the agent generated. For testers moving into SDET roles, this is a valuable skill: designing evidence that makes failures understandable and fixes verifiable. If you handed your CI failure report to an AI agent today, would it have evidence to investigate—or just a red badge? Reading that prompted this: z.ai/blog/glm-built-its-infe… #AI #TestAutomation #QualityEngineering #SDET
1
1
5
3,866
Your AI reports 0.95 confidence. That is not automatically a 95% chance of being right. TypeSafe's new article introduces Jev: a model built for structured decisions that software can use directly. One detail in its documentation deserves QA's attention: the confidence field summarises how concentrated the model's probability distribution is. It is not automatically your application's measured accuracy. My QA takeaway: test the decision policy, not just the score. Imagine an AI bug-triage tool confidently labelling a checkout failure “minor”. A valid label and a high score do not make the release safe. Here is how I would test it: 1. Define what the number means Document how the score is calculated and what it is intended to represent. Do not turn a 0–1 field into a percentage-correct badge without evidence. 2. Use cases the threshold was not tuned on Evaluate against a held-out, human-reviewed set of defects. Include missing context, unfamiliar inputs and rare critical failures—not only easy examples. 3. Measure mistakes AND automation coverage Of the cases handled automatically, how many were wrong? How many cases were sent for review? A threshold can appear safe simply because it escalates almost everything. 4. Test the boundary and the fallback Use controlled scores just below, at and above the threshold. Verify the correct route, a real review-queue entry and no unintended action when the score is missing or invalid. 5. Give critical errors their own release gate A strong average must not hide severe defects being downgraded. Keep deterministic safety rules outside the model and rerun the evaluation after model, prompt or policy changes. This is where test automation adds value: repeatable inputs, independent expected outcomes and evidence for when automation should stop. Does your AI test report measure confidence—or correctness at the point where the system acts? Reading that prompted this: typesafe.ai/blog/introducing… Confidence definition: docs.typesafe.ai/confidence #AI #TestAutomation #AITesting #QualityEngineering
2
1
8
12,666
Your voice AI can stop talking. But can it stop doing? Google’s new Gemini 3.8 Live announcement describes voice agents that keep a conversation going while tools and API calls run in the background. That creates a QA challenge: the conversation and the transaction can move at different speeds. Consider this test scenario: “Book the Friday appointment.” “Wait — make that Monday.” The agent says, “Of course.” But the Friday request is already in flight. A natural-sounding response is not proof that the booking is correct. Here is the automation coverage I would build: 1. INTERRUPT AT DIFFERENT STAGES Change the request before the API call, while it is pending, and after it commits. These need different expected outcomes. 2. VERIFY THE ACTUAL STATE Check the booking service, not just the transcript. An unsent action should stay unsent. An in-flight action needs reconciliation. A completed action needs an explicit change or cancellation flow — not a claim that it never happened. 3. CHALLENGE LATE RESULTS Return Friday’s result after Monday becomes the current request. Check that stale output cannot overwrite the newer intent or trigger another action. 4. TEST RECONNECTS AND RETRIES Drop the connection mid-task. Verify that recovery does not create duplicate appointments or report success without a confirmed result. 5. MEASURE TWO KINDS OF QUALITY Measure how quickly the agent stops speaking or acknowledges a correction. Separately measure whether the resulting action matches the user’s latest confirmed request. I would start with reviewed, repeatable scenarios in a sandbox, then vary speech timing, background noise and network delay. AI can help generate those variations. QA still needs to define the expected outcome. This is my proposed testing approach, not a claim that Gemini failed these scenarios. A smooth conversation is a UX win. A correct state transition is a QA requirement. What happens in your tests when the user changes their mind halfway through an action? Source inspiration — Google: blog.google/innovation-and-a… #AI #TestAutomation #QualityEngineering #VoiceAI
5
5
4,354
Your AI test agent says “100% passed”. Ask one more question: What did it never test? Vals AI’s recent write-up reports that Claude Fable 5.1 solved a centuries-old cipher. The experiment favoured problems with checkable answers. My QA lesson is different from the headline: An automation system must not make hard-to-verify risks disappear. Imagine this run: 100 checks planned 80 executed and passed 20 blocked or not run 100% pass rate among executed checks. 20 unresolved checks. Both facts belong in the release report. Here is the framework I would build: 1) Define the expected business outcome BEFORE AI generates the test. 2) Preserve separate passed, failed, skipped, blocked and not-run results. A timeout is not a pass. 3) Assert the real outcome. A refund banner does not prove the agreed amount or transaction state is correct. Check the API/state as well as the relevant UI behaviour in a safe test environment. 4) Reconcile planned critical checks with actual CI/CD results. Missing critical evidence stops automatic promotion and triggers review. 5) Use AI to diagnose gaps and draft improvements. Require reviewed assertions and a verified rerun before marking the gap resolved. This is a valuable manual-tester-to-SDET skill: Turn your knowledge of awkward business scenarios into checks that cannot quietly vanish from the dashboard. AI can accelerate execution. QA must keep the uncertainty visible. Does your pipeline report what was NOT verified? Source inspiration — Geby Jaff, Vals AI: vals.ai/blogs/fable-solves-c… #TestAutomation #AI #QualityEngineering
2
3
3,405
Your test code didn’t change. Your test result did. Before blaming “flaky automation”, check what changed underneath it. Homebrew 7.0.0 is out today. Its release notes cover changes to platform support, CI workflows and built-in vulnerability checks. For QA teams, that raises a bigger question: Are you versioning your test scripts—or the conditions that make their results meaningful? Imagine the same UI/API suite running on two machines. One uses yesterday’s browser and dependencies. The other installs newer versions during setup. One fails. The application may have regressed. The environment may have changed. Without evidence, you are guessing. Here is the environment contract I would add to a test framework: 1. RECORD WHAT ACTUALLY RAN Attach the application commit, OS, CPU architecture, runtime, browser and dependency versions to the test report. Record the container image digest where relevant. Never publish secrets in diagnostics. 2. PROVE A CLEAN START Run a scheduled build on a fresh runner with no inherited developer tools. Install declared dependencies, seed synthetic data and execute a small UI/API smoke pack. A failed setup must be reported as a failed setup—not a successful test run. 3. TEST THE SUPPORTED MATRIX Choose environments from product and team requirements. Check the combinations you actually support, including relevant OS and architecture differences. Identical scripts do not prove identical coverage. 4. SEPARATE COLD AND WARM RUNS Test with an empty cache and with the expected cache restored. Check both execution results and timing. A fast pipeline that depends on an undocumented cache is difficult to reproduce. 5. REVIEW TOOLCHAIN CHANGES LIKE CODE Use locked dependencies and immutable references where supported. Trial upgrades in a reviewable change, compare results and keep a recovery path. Controlled updates—not frozen, unpatched dependencies. Where does AI fit? Use it to compare sanitised environment manifests, explain setup errors and propose migration changes. Review those changes before adoption. Do not accept weakened assertions or disabled security controls merely to make CI green. Homebrew’s new vulnerability checks can contribute evidence. They do not certify your entire test environment as secure. For a manual tester becoming an SDET, this is a valuable shift: learn to explain why a test result is trustworthy, not only how to automate a click. Could someone rebuild your test environment tomorrow and explain every version difference? Source: Homebrew 7.0.0 release notes brew.sh/2026/09/13/homebrew-… #TestAutomation #CICD #DevOps #AI
5
2
3,176
A fast API that saves the wrong data is not a successful release. Database scaling is a QA story—not just an infrastructure story. PlanetScale has announced Neki: Postgres distributed across multiple machines, with built-in workflows for operations such as resharding and failover. Important: it is a platform preview. PlanetScale explicitly says not to run production workloads on it yet. The announcement raises a bigger testing question: When the infrastructure changes underneath your application, which business promises must remain true? Sharding means splitting data across database partitions. Your users should not need to understand that to trust an order, booking or account balance. Here is the QA automation strategy I would build around a sharded application: 1. TEST BUSINESS RULES, NOT JUST RESPONSE CODES For an order service, an accepted retry must not create a second order. A customer must never receive another customer's records. Verify stored outcomes and access rules—not just a 200 OK. 2. TEST SEQUENCES, NOT ONLY SINGLE REQUESTS Create an order. Change it. Retry after a timeout. Submit two competing updates. Read it back. Check each outcome against an independently reviewed business model and the documented transaction guarantees. 3. TEST WHILE THE SYSTEM CHANGES In an isolated test environment, run supported maintenance and recovery scenarios while representative traffic continues. Compare records and business totals before, during and after. Check recovery time as well as correctness. 4. TEST UNEVEN DEMAND Do not distribute every test customer evenly. Model one very busy tenant alongside many quiet ones. Measure tail latency, errors and correctness per tenant—not just a reassuring global average. These are proposed application tests, not claims that Neki has these faults. Where does AI help? Use it to propose synthetic datasets, generate workflow variations and cluster failure evidence. Keep the expected results human-reviewed, retain reproducible test inputs, and keep customer data out of unapproved AI tools. AI can help create the workload. It should not invent what “correct” means. For manual testers becoming SDETs, this is a valuable step beyond scripting clicks: turn your knowledge of business rules into automated evidence that survives infrastructure change. If your database doubled its capacity tomorrow, could your tests prove the business still works—not merely that the API is faster? Source: PlanetScale Engineering planetscale.com/blog/introdu… #TestAutomation #PostgreSQL #AI #QualityEngineering
2
1
1,979
AI can write your mobile app twice. Your QA strategy still has to prove it works on both platforms. Shopify has announced a move from React Native back to Swift and Kotlin, saying coding agents have changed the economics of building for iOS and Android. But the most useful QA detail is not the framework switch. Shopify describes separating business logic from the UI and exposing it through a CLI, so agents can test that logic without repeatedly driving a simulator. Its migration process also uses small checkpoints, tests, visual reviews and human approval. My takeaway: make the product easier to test before asking AI to generate more tests. Here is how I would structure the automation: 1. FAST BUSINESS-RULE CHECKS Test calculations, validation and state transitions without opening a screen. A discount rule should be verifiable without tapping through checkout 100 times. 2. SHARED API AND DATA CONTRACTS Give iOS and Android the same reviewed business examples and expected outcomes. Verify the API response and persisted result—not just a success message. 3. PLATFORM-SPECIFIC PROOF Keep real UI/device checks for what headless tests cannot establish: permissions, keyboards, deep links, accessibility, background/resume behaviour and rendering. A testable core does not make the UI irrelevant. It lets each test run at the layer where it provides useful evidence. Use AI to propose edge cases, draft automation and analyse failures. Keep expected business outcomes independently reviewed. Run fast checks on each change, then the relevant integration and device checks before release. This is not a claim that every team should abandon React Native. It is a reason to examine how much your architecture helps—or obstructs—testing. For manual testers moving towards SDET work, this is a valuable skill: translating product knowledge into checks at the right layer, not simply converting every manual step into a click script. If your app’s UI disappeared tomorrow, how much of its business behaviour could your test suite still verify? Source: Shopify Engineering shopify.engineering/back-to-… #AI #TestAutomation #MobileTesting #QualityEngineering
2
2
7
7,789
Same model name. Smaller download. That is still a QA release decision. Quesma’s recent Qwen3.8 benchmark found that a 4-bit version held up on the tested coding benchmark, while a 1-bit version deteriorated sharply on other tasks. Quantisation means storing a model’s numbers with fewer bits to reduce its memory footprint. Those results are specific to that study—not a guarantee for your application. My QA takeaway: the model name is not the release configuration. A team can keep the same UI, API and prompts while changing the model underneath. Yesterday’s test report may no longer describe what customers are using. Here is how I would automate that release check: 1. Give the deployed model a fingerprint Record the exact model artifact and checksum, runtime version, compression format, prompt version and generation settings. For a hosted model, retain the version and configuration the provider exposes. Attach these to every test report. 2. Test the real path—not only the mock Keep fast mocked checks, but run a bounded integration suite against the candidate configuration. Validate structured API outputs and a few critical UI journeys. A green test against a stub says nothing about the replacement model. 3. Score business failures separately Imagine an AI support tool that returns valid JSON but selects the wrong customer record. Schema validation passes; the workflow is wrong. Use synthetic records, explicit expected outcomes and independent permission checks. Test ambiguous requests, missing information and long inputs too. 4. Measure useful work, not cheap tokens Repeat important evaluations. Report task success, unsafe actions, latency, retries and cost per correctly completed task. Agree risk-based release thresholds before seeing the results; keep a tested rollback route. For an AI-assisted test-writing tool, add another check: do its generated tests detect seeded defects, or merely pass against the code it just wrote? My view after 20 years in test automation: a model configuration change deserves the same release discipline as a code change. This is valuable SDET work. Product knowledge becomes test scenarios, API assertions, CI/CD gates and evidence people can trust. If you are moving from manual testing into SDET, or upgrading your team’s automation, book a free 30-minute conversation: calendly.com/mitchellagoma/f… Could your team reproduce the exact AI configuration behind its last green test report? Reading that prompted this — Piotr Migdał, Quesma: quesma.com/blog/qwen38-27b-q… #TestAutomation #AI #SDET
2
2
2,319
Your AI refactor is faster. What else changed? A browser engineering update offers a useful QA lesson for the AI era: compare behaviour before celebrating speed. In its August update, Ladybird describes checking a new style engine against a simple reference evaluator, and checking cached layout results against fresh calculations. My takeaway: use differential testing—give two implementations the same inputs and investigate the differences. CONTROLLED INPUTS ↓ REFERENCE + NEW IMPLEMENTATION ↓ COMPARE RESULTS AND SIDE EFFECTS ↓ EXPLAIN EVERY MEANINGFUL DIFFERENCE Here is how I would apply that to an AI-assisted refactor: 1. Define what must stay the same. Business rules, permissions, errors and persisted state. Separate intentional behaviour changes from performance work. 2. Build representative inputs. Include realistic records, boundary values and sequences of actions—not just random payloads that fail validation immediately. 3. Compare more than HTTP 200. Check response data, database changes and emitted events. Normalise only genuinely irrelevant differences, such as generated trace IDs. 4. Keep the comparison safe. Use isolated test environments or side-effect-free shadow execution. Never send the same payment or email twice just to compare implementations. 5. Treat mismatches as evidence. AI can help minimise a failing input and group similar differences. A person still needs to decide whether the new result is a defect, an intended change or a correction to an old bug. Imagine AI adds caching to a pricing API. The endpoint becomes faster. But after a customer changes quantity, the cached total stays the same. A latency chart celebrates. A sequence-aware comparison catches the problem. Two implementations agreeing is not proof of correctness; both may share a bug. Keep independent business-rule checks too. For manual testers moving into SDET, this is valuable work: deciding what to compare, which differences matter and how to make the comparison repeatable in CI. Don't just ask AI to optimise the code. Ask your test strategy to expose what the optimisation changed. Which refactor would you test this way first? Reading that prompted this: ladybird.org/newsletter/2026… #TestAutomation #AI #QualityEngineering #SoftwareTesting
2
3
1,285
Your AI agent found a bug. Will CI catch it when it comes back? AI exploration is a discovery engine—not permanent regression protection. The missing step is turning a useful finding into a check your team can trust next week. DISCOVER → REPRODUCE → REVIEW → AUTOMATE → RUN IN CI Here’s the practical version: 1. Bound the exploration. Choose the feature, customer risk and safe environment. Limit access, actions and spend. 2. Capture the evidence. Keep the steps, test data, build version and traces. “The agent says it failed” is a lead, not a verified defect. 3. Confirm what should happen. Reproduce the failure. Check the business rule. Rule out bad data and environment problems. 4. Build the right regression check. Use an API check for the business rule; add UI coverage where the interaction matters. 5. Prove the check works. It must fail against the defect and pass after the fix. Keep it in CI with controlled setup and a named owner. Example: a read-only user can rename a workspace through the API, even though the UI hides the edit button. Checking that the button is hidden is not enough. Assert that the API rejects the change AND the workspace name stays unchanged. Keep exploration for discovery. Keep explicit regression checks for repeatable protection. Keep people accountable for the risk. After 20 years in test automation, this is the shift I want more teams to make: from collecting bug reports to preserving what they learn. Manual tester moving into SDET—or a QA team adopting automation? I can help with API/UI frameworks, CI/CD and AI-assisted testing. Book a free 30-minute strategy call: calendly.com/mitchellagoma/f… How many bugs from your last exploratory session now have a reliable regression check? Inspired by antirez’s “A new era for software testing”: antirez.com/news/168 #AI #TestAutomation #SDET #Playwright
2
2
1,914