0
I’m looking at Claude’s computer-use capabilities for repetitive browser tasks, but I’m unsure how to judge whether an action is based on the current screen rather than an outdated assumption about where a button or dialog should be.
For people testing this, what kinds of checks help catch a click on the wrong page or a task that quietly went off track? I’m interested in practical evaluation approaches, especially when a workflow has several steps and the interface can change as it runs.