This is very interesting! I saw a lot of variance in task outcomes for frontier models in our experiments on computer use and robotics tasks too. This really makes you wonder where eval should go beyond the benchmarks today.
btw, we just put out the full per-task breakdown, in case it's useful. Here's our report on it too. Would love your thoughts