Welcome.AIWelcome.AI
    Skip to content
    research

    StoreBench: Evaluating Autonomous Agents in Live-Commerce Settings

    Recent research has introduced StoreBench, a dynamic testing environment designed to enhance the capabilities of large language models (LLMs) in post-training scenarios. Unlike traditional benchmarks...

    arxiv.org•October 9, 2026•4 min read

    Key Facts

    • Implement StoreBench to accurately evaluate LLM performance in dynamic retail environments.
    • Utilize insights from LLM interactions to enhance decision-making processes in e-commerce.
    • Adjust marketing strategies based on real-time customer demand simulations provided by StoreBench.
    • Leverage StoreBench's consistent testing framework to improve model training and performance reliability.
    • Analyze supplier unpredictability to develop adaptive supply chain strategies using LLM insights.

    Summary

    Paper: StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents

    Authors: Daksh Raghuvanshi, Ved Vedere, Yifan Wang

    Executive Summary

    Recent research has introduced StoreBench, a dynamic testing environment designed to enhance the capabilities of large language models (LLMs) in post-training scenarios. Unlike traditional benchmarks that rely on static evaluations, StoreBench simulates a live-commerce setting where an agent operates an online apparel store. This environment reflects real-world complexities, such as fluctuating customer demands, supplier unpredictability, and sudden market changes, providing a more realistic backdrop for evaluating LLM performance.

    In StoreBench, the agent interacts with the same 29 tools that human merchants use, allowing for a genuine comparison of decision-making processes. The environment is structured to maintain consistency, with each episode replaying identically based on the sequence of actions taken. This setup ensures that results are not influenced by variations in model latency, making it a more reliable testing ground.

    The research evaluated seven leading LLMs across 11 different scenarios that simulated timeframes ranging from 30 to 45 days, as well as a comprehensive full-year simulation. The findings reveal that no model could match the performance of a scripted smart-triage policy, which achieved an average success rate of 97%. The best-performing model in the study, DeepSeek-V4-Pro, managed to pass only 49% of the task-seed cells, highlighting a significant gap between machine and human performance. In fact, human experts using the same tools outperformed all models, achieving a mean composite score of 0.708 compared to the models’ 0.700.

    Interestingly, the study also noted that performance improved significantly for several models over the course of the simulated year, particularly in a specific post-training run involving the Qwen3.5-27B model. This model, initially trained on five different tasks, saw its average score rise from 0.136 to 0.373, indicating that targeted training can enhance model capabilities even in complex environments.

    The research contributes valuable insights into the effectiveness of LLMs in real-world applications, particularly in sectors like e-commerce where decision-making under uncertainty is critical. The introduction of StoreBench as a rigorous testing environment may help future developments in AI by providing a clear benchmark for evaluating LLMs' operational capabilities. The authors have released several example training tasks and verification tools to support further research, although the complete evaluation suite remains restricted to prevent contamination of the benchmark results.

    Overall, this study emphasizes the necessity for more dynamic and realistic benchmarks in AI development, particularly for applications that require advanced planning and economic reasoning. The insights gained could inform future advancements in AI technologies, potentially benefiting industries that rely on complex decision-making processes.

    Academic Abstract

    Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use, under a windowed operation budget that makes simulated time a function of actions taken, so model latency cannot influence simulated time. Pass thresholds are calibrated against scripted anchor policies, the reward is hardened against a catalogue of reward hacks, and every episode replays identically given a sequence of actions. We evaluate seven frontier LLMs on 11 scenarios of 30 to 45 days and a full simulated year, over three world seeds at matched reasoning effort. No model matches the scripted smart-triage policy on average: the best, DeepSeek-V4-Pro, passes 49% of task-seed cells against the heuristic's 97%. Human experts working through the same tools and budgets outscore every model (mean composite 0.708 vs. 0.700). Over a full simulated year under the Claude Code harness, most models show dramatic performance improvement. In a GRPO post-training run, Qwen3.5-27B trained on only five disjoint tasks raises its mean composite on the held-out evaluation tasks from 0.136 to 0.373. We release five example training-split tasks, ten sample trajectories, and the scoring and verification tooling; the full environment and evaluation suite are withheld to keep the benchmark uncontaminated.

    Frequently Asked Questions

    What business problems does StoreBench address?

    StoreBench addresses the challenges of evaluating and training autonomous operator agents in a dynamic live-commerce environment, specifically dealing with fluctuating customer demands, supplier unpredictability, and sudden market changes.

    Which industries may benefit most from the findings of StoreBench?

    The retail and e-commerce industries may benefit most, particularly those focusing on online apparel sales, as StoreBench is designed to simulate an online store environment relevant to these sectors.

    What are the practical implementation considerations for businesses using StoreBench?

    Businesses implementing StoreBench should consider the need for a robust simulation framework that accurately reflects live-commerce dynamics and the necessity of integrating the same tools used by human merchants for effective agent training.

    What resources or expertise are needed to utilize StoreBench effectively?

    Organizations may need access to advanced large language models (LLMs), expertise in machine learning and AI, as well as resources to develop and maintain the dynamic testing environment required for conducting evaluations.

    What competitive advantages could businesses gain from using StoreBench?

    By utilizing StoreBench, businesses could gain competitive advantages through improved decision-making processes of autonomous agents, leading to better responsiveness to market changes and enhanced customer satisfaction in their online operations.

    Where AI Leaders Stay Informed

    The latest AI intelligence, case studies, and research — delivered to your inbox every week.

    Free to read. Unsubscribe anytime.