Limitations of Large Language Models Exposed: A Benchmarking Study on Web Agent Performance

Wednesday 09 April 2025


The quest for a web agent that can truly assist us in our daily online lives has been ongoing for some time now. Researchers have been working tirelessly to develop systems that can navigate the vast expanse of the internet, retrieve relevant information and provide answers to our queries. Recently, a team has made significant strides in this direction, introducing BEARCUBS, a benchmark designed to evaluate the capabilities of web agents.


BEARCUBS is not just another set of questions and answers; it’s a comprehensive evaluation framework that assesses an agent’s ability to interact with real-world websites, search for information, and provide accurate responses. The team behind BEARCUBS has crafted 111 information-seeking questions that require agents to demonstrate their skills in various areas, such as identifying factual information, navigating complex web pages, and understanding multimodal interactions.


One of the key challenges faced by current web agents is their inability to handle dynamic and unpredictable real-world web interactions. BEARCUBS addresses this issue by incorporating realistic scenarios into its evaluation framework. The questions are designed to mimic how humans would interact with websites, making it a more accurate representation of real-world usage.


To evaluate the performance of various web agents, the researchers created a baseline setting using two popular language models, gpt-4o and DeepSeek R1. They also included human performance as a benchmark to gauge the effectiveness of each agent. The results are revealing: while some agents performed reasonably well on text-based questions, their abilities to navigate multimodal interactions were severely lacking.


Grok 3, one of the language models used in the baseline setting, was found to return its search and reasoning trajectory in non-English languages if present in the input, making it less helpful for users seeking assistance in those languages. Another notable issue is that Grok 3 never abstains from providing an answer, even when faced with impossible tasks.


The study highlights the need for improved adaptability, transparency, and user-centered refinement in LLM-based computer use agents. The researchers emphasize that these issues are not unique to BEARCUBS but rather reflect broader challenges faced by web agents as a whole.


As we continue to rely more heavily on technology to navigate our online lives, it’s essential that we develop systems that can truly assist us. BEARCUBS is an important step in this direction, providing a comprehensive evaluation framework for web agents.


Cite this article: “Limitations of Large Language Models Exposed: A Benchmarking Study on Web Agent Performance”, The Science Archive, 2025.


Web_Agents, Bearcubs, Benchmark, Language Models, Gpt-4O, Deepseek R1, Multimodal Interactions, User-Centered Refinement, Adaptability, Transparency, Llm-Based Computer Use Agents.


Reference: Yixiao Song, Katherine Thai, Chau Minh Pham, Yapei Chang, Mazin Nadaf, Mohit Iyyer, “BEARCUBS: A benchmark for computer-using web agents” (2025).


Leave a Reply