Last Updated on 7. October 2026
What if an AI agent could not only write code but also verify its functionality? With QF-Test 11, that is no longer a preview scenario. The new version of the UI test automation tool from the mgm group brings agentic testing into everyday QA work in two directions: AI works inside QF-Test, and external AI agents use QF-Test as a tool.
From Recording to Instruction: How Test Automation Is Fundamentally Changing
Anyone who develops software knows the problem: every update, every new feature, and every code change needs to be verified. Did everything still work the way it did before? Is the application displaying what it’s supposed to display? Traditional test automation has already simplified this process considerably: a tool takes over the clicking, checks the results, and runs through everything automatically again and again, reliably, repeatably, and fast. QF-Test has been used for exactly this purpose for years, covering web and desktop applications, Java GUIs, mobile apps, and PDF documents.
But now the game is changing again. Because the same development teams that today use Claude Code, Cursor, or GitHub Copilot to write code also want to use those same agents for testing, without ever having to open the test tool. “Our customers are incredibly interested in figuring out which parts of the testing process they can hand off to AI,” explains Max Melzer, software developer and trainer on the QF-Test team. This comes as no surprise: anyone who has watched an AI agent independently refactor code will inevitably ask why that same agent can’t handle testing too.
QF-Test 11 answers this question with two sets of features. The first brings AI into QF-Test. The second opens QF-Test up to AI agents.
AI Inside QF-Test: Generating and Running Tests in Natural Language
QF-Test takes a staged approach — from initial AI-assisted checking capabilities all the way to full integration into agentic development pipelines. All three levels build on one another and are part of ongoing product development.
Test suite generation: a first draft in minutes
Instead of recording or writing tests manually, AI handles the first draft. In QF-Test 11, an assistant for this sits right in the toolbar. Testers describe what they want to test, and the integrated AI agent turns it into a structured draft. QF-Test can also explore the application independently. The agent examines the running application live, identifies the most important functions and paths, and proposes test cases. Alternatively, it works from existing documentation.
For example, a prompt in a Swing application asks for tests of a single discount button and sets clear limits: stay on the same tab and don’t open any dialogs from the menu. Along with the test steps, QF-Test delivers a test description and a test plan. For version 11.1, the plan is to convert as many steps as possible into regular QF-Test nodes. AI will then only serve as a fallback for the rest. This makes the results reproducible, because a regular node behaves the same way in every test run.
“It’s never the case that you can just take it and say: we’re done now,” explains Max Melzer. “But it takes a lot of work off your plate — especially at the start.” A solid first draft for experienced testers to build on. Why this distinction matters is the subject of our article Between AI Euphoria and Quality Engineering.
AI instructions: plain language as a test step
Generated tests contain the second new building block, the AI Instruction node. In this test step, testers tell the AI in plain language what to do or verify, for example “Open the application’s Info dialog.” The selected model then uses QF-Test’s connection to the application to carry out the instruction. Under the hood, it relies on internal MCP tools. The QF-Test run log records which tools the AI called and what happened. This keeps the step traceable even though nobody defined it in detail.
AI checks and test data
Some results can’t be captured with simple pass/fail checks. A prime example is testing integrated chatbots. Their responses are never word-for-word predictable, and yet the test still needs to verify whether the response is correct in substance. Since QF-Test 10, a dedicated node passes the result to a language model for evaluation. Is this response plausible? Does it fall within the expected parameters? The model makes a decision, and QF-Test incorporates the result into the test run. Semantic testing instead of rigid string comparisons.
New in QF-Test 11 is AI-generated test data. Testers can fill data tables with matching values in just a few clicks. The prompts behind the generation can be inspected and adjusted, for instance to get French names and addresses wherever possible.
Bring your own model
For all of its AI features, QF-Test follows a “bring your own model” approach. Teams decide for themselves which provider and which model get access to their test data. Organizations with strict data handling requirements use internal or self-hosted models. A smaller model also keeps costs down. According to the QF-Test team, the tools are built so that smaller models work well with them too. Claude Code and GitHub Copilot are now officially supported as command-line tools. And where AI is not an option at all, the AI features can be switched off entirely, a setting available since QF-Test 10.1.
QF-Test as a Tool for AI Agents: MCP Server and MCP Toolkit
The second direction is the heart of the agentic approach: with QF-Test 11, the tool gains an integrated MCP server. MCP (Model Context Protocol) is an open standard that allows AI agents to access external tools in a standardized way.
In practice, this means: anyone using Claude Code, GitHub Copilot, or another AI coding agent in their development environment can give that agent access to QF-Test’s capabilities. The agent can then, without ever manually opening QF-Test, launch applications, control websites, desktop applications, and apps, run tests, check results, and report back errors. QF-Test becomes part of the agentic workflow: a tool the AI agent calls just like any other. The MCP server runs both interactively and in batch mode, which also makes it suitable for automated pipelines.
QF-Test’s MCP toolkit: making any application AI-ready
QF-Test 11 goes one step further. With the MCP toolkit, teams define their own MCP tools: they record a sequence of steps in QF-Test, turn it into a procedure, describe it with parameters and return values, and activate it as an MCP tool. QF-Test then acts as a kind of proxy, making applications controllable by AI even though they were never built for it.
This brings several advantages. For one, existing testing know-how stays useful, because recorded sequences and existing procedures become tools that an AI agent can call on purpose. Instead of clicking through the user interface step by step, the agent simply picks the right tool. This saves tokens and makes execution faster and more predictable. The results remain verifiable. If an outcome doesn’t match the expectation, QF-Test detects the error and logs it just as it would in any other test. Since version 11.0.2, an overview page in the options dialog lists all available MCP tools.
QF-Test’s Edge in Agent-Based Testing: Breadth Beats Niche
What sets QF-Test apart in this space isn’t MCP capability alone. Specialized tools like Playwright now offer that too. The decisive difference lies in platform breadth: unlike many competitors that focus exclusively on web applications, QF-Test tests native Windows applications, Java GUIs, Android and iOS apps, and PDF documents using the same approach. Enterprises and public authorities rarely run a uniform application landscape. With QF-Test, they get a unified test agent for all of it. And with the MCP toolkit, even older specialist applications can become part of an agentic workflow.
| Platform | QF-Test | Playwright |
|---|---|---|
| Web Applications | ✓ | ✓ |
| Native Windows Apps | ✓ | – |
| Java GUIs (Swing, SWT, JavaFX) | ✓ | – |
| Native Android & iOS Apps | ✓ | – |
| PDF Documents | ✓ | – |
More in QF-Test 11
Not everything in the new version is about AI. During recording, QF-Test now detects matching procedures from your own library and inserts the calls automatically. The same mechanism can tidy up existing test cases. A new OpenAPI import creates a complete test suite from an API specification in seconds, including full create, read, update, and delete sequences for each resource. And a command palette gives keyboard access to almost every function and to all elements of your test suites, with a search that also understands fuzzy terms.
Where Things Stand Today: Webinar and Roadmap
The best way to see the new features in action is the webinar “AI agents in action: autonomous test generation and MCP with QF-Test 11” from August 2026. It covers the AI configuration, test data and test suite generation, AI instructions, and examples for the MCP toolkit. Slides and demo suites are available for download on the webinar page.
For a compact overview of all new features, watch the QF-Test 11 highlights video.
The roadmap for the rest of 2026 shows where development is heading: advanced security features for agentic actions and AI-generated test suites with more fixed, non-dynamic steps. Both aim to give the agent more autonomy while keeping test results reproducible.
QF-Test 11 can be downloaded and tried out at any time. Saving test suites requires a license, and a free trial license is available.
Would you like to see what agentic testing with QF-Test 11 can do for your application landscape? Download the current version and request your free trial license here: https://services.qftest.com/en/license/request/
FAQ: Agentic Testing with QF-Test
What is agentic testing?
Agentic testing is an approach to software quality assurance in which artificial intelligence (AI) agents independently plan, execute, and evaluate test tasks. QF-Test relies on firmly defined test building blocks: starting with version 11.1, generated steps are set to become regular QF-Test nodes wherever possible, with AI only stepping in where it is needed. This keeps test runs reproducible and traceable. QF-Test supports agentic testing for web applications, native Windows applications, Java GUIs, Android and iOS apps, and PDF documents.
What AI features does QF-Test 11 include?
QF-Test 11 generates test suites from a description or by exploring the application, executes AI instructions written in plain language, and generates test data. AI checks for non-deterministic results such as chatbot responses have been available since QF-Test 10. In addition, a built-in MCP server and the MCP toolkit make QF-Test usable for external AI agents.
What is an MCP server in the context of test automation?
MCP stands for Model Context Protocol, an open standard that facilitates communication between AI agents and external tools. An MCP server makes a tool’s capabilities (such as those of QF-Test) available to AI agents, such as Claude Code or GitHub Copilot. With QF-Test’s MCP server, agents launch applications, control user interfaces, and run tests. With the MCP toolkit, teams turn their own QF-Test procedures into MCP tools.
Which AI models can I use with QF-Test?
QF-Test follows a “bring your own model” approach. You configure the provider and model yourself, including internal or self-hosted models. Claude Code and GitHub Copilot are officially supported as command-line tools. If required, all AI features can be switched off entirely.
Does agentic testing replace human testers?
No, AI-generated test suites provide a structured first draft that experienced testers must further develop and validate. The value lies in reducing the effort required to create and maintain tests, not in replacing human expertise entirely.





