Work done with Katherine Thai
@chautmpham.bsky.social @yapeichang.bsky.social Mazin Nadaf & @miyyer.bsky.social
🌐 bear-cubs.github.io/
Work done with Katherine Thai
@chautmpham.bsky.social @yapeichang.bsky.social Mazin Nadaf & @miyyer.bsky.social
🌐 bear-cubs.github.io/
💡 Stronger multimodal reasoning (videos, maps, real-time data)
🔍 More reliable source selection
🗺️ Smarter and more efficient search strategies
📜 Transparent and interpretable browsing trajectories
💡 Stronger multimodal reasoning (videos, maps, real-time data)
🔍 More reliable source selection
🗺️ Smarter and more efficient search strategies
📜 Transparent and interpretable browsing trajectories
Current agents struggle with:
🚨 Selecting reliable sources
🚨 Escaping dead loops
🚨 Engaging in multimodal interactions
🚨 Navigating the web in real-time
Current agents struggle with:
🚨 Selecting reliable sources
🚨 Escaping dead loops
🚨 Engaging in multimodal interactions
🚨 Navigating the web in real-time
Not great ...
🥴 The best multimodal web agent, OpenAI’s Operator, scores 24.3% accuracy.
🤯 OpenAI’s Deep Research outperforms all (35.1%), without computer-use abilities!
Not great ...
🥴 The best multimodal web agent, OpenAI’s Operator, scores 24.3% accuracy.
🤯 OpenAI’s Deep Research outperforms all (35.1%), without computer-use abilities!
1️⃣ Use simulations (e.g., WebArena), missing real-world complexity
2️⃣ Have limited multimodal testing, relying on HTML (Mind2Web) or specific skills (e.g., map)
3️⃣ Are nearing performance saturation—Operator hits 87% on WebVoyager
1️⃣ Use simulations (e.g., WebArena), missing real-world complexity
2️⃣ Have limited multimodal testing, relying on HTML (Mind2Web) or specific skills (e.g., map)
3️⃣ Are nearing performance saturation—Operator hits 87% on WebVoyager
🔹Benchmarks computer-using agents @OpenAI Operator, @AnthropicAI Computer Use, and @convergence_ai_ Proxy
🔹Evaluates complex text-based & multimodal interactions
🔹Will be updated regularly with new questions
📜 arxiv.org/abs/2503.07919
🌐 bear-cubs.github.io/
🔹Benchmarks computer-using agents @OpenAI Operator, @AnthropicAI Computer Use, and @convergence_ai_ Proxy
🔹Evaluates complex text-based & multimodal interactions
🔹Will be updated regularly with new questions
📜 arxiv.org/abs/2503.07919
🌐 bear-cubs.github.io/