
Google Cloud AI Research and academic partners have developed an open source framework called EnvHarness, which allows AI agents to train against environments that evolve with them. This framework turns static environments into dynamic ones, enabling agents to learn and improve more effectively.
Training agents for specific tasks requires environments where they can practice, fail, and improve. However, building these environments is expensive, and once created, they typically remain static. EnvHarness addresses this issue by adding a programmable layer around an existing environment, allowing it to change and adapt to the agent’s weaknesses.
How EnvHarness Works
EnvHarness provides three components: Stage, Contract, and Chain. Stage changes the environment’s starting state, Contract modifies the interaction between the agent and environment, and Chain joins tasks together to create more complex scenarios. These components enable EnvHarness to reshape the training experience without modifying the underlying environment or its verifier.
The framework also includes EnvRigger, which automates the process of deciding how EnvHarness should modify the environment based on the agent’s weaknesses. EnvRigger follows an “Observe → Diagnose → Write → Validate” loop, examining the agent’s trajectories, composing EnvHarness components, and validating the changes to ensure they create useful and solvable training examples.
In experiments, EnvHarness was tested on five benchmarks, including software engineering, web navigation, and office work. The results showed that agents learning from EnvHarness environments improved by up to 9 points on held-out tasks and accomplished tasks in fewer steps than agents learning from the original environments.
Benefits for Enterprise AI Teams
EnvHarness presents an alternative to continuously building new simulators and training tasks from scratch. By starting with a trusted environment and dynamically reshaping it around the agent’s weaknesses, enterprise AI teams can reduce development overhead and get more coverage from each environment. EnvHarness can be paired with skill or memory extraction, fine-tuning, reinforcement learning, or other mechanisms to change the agent based on its experience.
Related Post: Google’s Dream-RSI slashes discovery-agent calls by 162x
According to Zifeng Wang, Research Scientist at Google and co-author of the paper, EnvHarness does not replace the underlying environments but acts as an amplifier for them. “While human-designed environments will continue to serve as the baseline ground truth, approaches like EnvHarness may offer a practical path to reducing development overhead and getting more coverage from each environment,” he said. The researchers have released the EnvHarness code, experiment configurations, and RL implementation on GitHub under the permissive Apache 2.0 license.
One of the key advantages of EnvHarness is its ability to create dynamic environmental constraints that can force an improved agent harness to confront behaviors it might otherwise avoid or solve through brittle shortcuts. This approach can be complementary to frameworks that automatically modify the agent harness, such as Self-Harness, HarnessX, and DarwinX, and could potentially form a feedback loop with these systems.
The implementation costs of EnvHarness include integrating an existing environment with the framework and the computational expense of the EnvRigger loop. However, the researchers expect the cost of designing and validating environment modifications to fall as the models that power EnvRigger improve. For now, EnvHarness is best suited for digital sandboxes where rollouts are cheap and state can be restored quickly, such as coding environments and web automation test systems.
EnvHarness has been tested on various benchmarks, including ALFWorld, WebArena, SWE-bench Verified, OfficeQA, and SpreadsheetBench. The results show that agents learning from EnvHarness environments outperformed those learning from unchanged environments across all five benchmarks.
In addition to improved accuracy, training with EnvHarness resulted in shorter average trajectories. On SWE-bench Verified, the average trajectory shortened from 55.01 to 49.61 steps. EnvHarness also compared favorably with systems built specifically to generate training environments, exceeding SWE-smith by 2.46 percentage points and GenEnv by 5.7 points.
Implementation and Future Directions
As the models that power EnvRigger improve, the goal is to reduce development overhead and get more coverage from each environment, rather than replacing human-designed environments. EnvHarness acts as an amplifier for existing environments, offering a practical path to improving agent training and performance.


