Please turn JavaScript on
Association for Software Testing icon

Association for Software Testing

Is this your feed? Claim it!

Publisher:  Unclaimed!
Message frequency:  1.12 / day

Message History

In this episode, we examine the expansion of Bedrock AgentCore Evaluations to generic frameworks using OpenTelemetry GenAI semantic conventions. We also analyze Escape Tech’s penetration testing methodology for Model Context Protocol servers, runtime lifecycle hooks for tool guardrails in Strands Agents SDK, and empirical findings on managing the handoff tax in multi-model ag...


Read full story

Piping full agent execution traces directly into an LLM evaluator often causes the judge to inherit generation bias, leading to unearned passing scores. A context-isolated evaluation architecture separates the grading evidence from the generation path, relying on deterministic assertions for structural criteria and calibrated model scoring for semantic analysis. This approach...


Read full story

Catchpoint has officially ended support for Selenium Transaction Tests, making migration to Playwright or Puppeteer mandatory for synthetic monitoring suites. Cypress 15.21.1 delivers reliability improvements for live agent debugging via cypress tap. In addition, new research highlights Inertia Bias, showing why AI evaluators should not share context with the planning steps t...


Read full story

Iterative repair of AI-generated tests can silently degrade assertions, prioritizing execution success over actual verification strength. This episode breaks down how to escape the self-repair trap by separating operational tool recovery from oracle repair. Discover practical methods like dual-context grounding and independent resampling to ensure your automated tests continu...


Read full story


Read full story