The focus of competition among AI agents is shifting from computer manipulation capabilities to operational systems that can be safely integrated into corporate tasks. It is analyzed that not only model performance but also context management, authority management, error recovery, and validation systems tailored to individual companies must be in place to entrust actual work to these agents.
Andreessen Horowitz stated in a post published on August 28, 2025, that computer-using agents are directly handling browsers and desktop environments, thereby expanding the scope of work automation. In corporate settings, orchestration that coordinates company-specific context and reliability design is deemed more important than the universal model itself.
Computer-using agents are AI systems that recognize the screens used by humans and perform actions such as clicking, inputting, and navigating. Unlike traditional robotic process automation (RPA) that repeats predefined procedures, these agents select necessary tools and process multiple steps after receiving a goal.
The application targets tasks that involve switching between various programs, such as document searches, customer management system inputs, internal messenger notifications, and report writing. Corporate tasks are intertwined with access permissions, exception handling, audit trails, and human approval processes, making it difficult to automate the entire process with screen manipulation capabilities alone.
According to the official OSWorld paper, the human performance rate in 369 actual computer tasks was 72.36%, while the highest model success rate at that time was 12.24%. OSWorld is a benchmark that measures the ability to handle files, applications, and web pages in real operating system environments.
OpenAI reported on January 23, 2025, that its Computer-Using Agent (CUA) achieved 38.1% in OSWorld, 58.1% in WebArena, and 87% in WebVoyager. These figures indicate that models directly handling computers have moved from the research stage to the productization stage.
However, even when measuring the same computer usage capabilities, it is difficult to make simple comparisons if the benchmark versions, task configurations, and execution harnesses differ. The original paper and OSWorld-Verified evaluations from individual companies apply different conditions, so performance figures must be accompanied by the evaluation criteria.
This is also why it is challenging to explain the bottleneck of enterprise AI agents solely through model scores. Even if a strong model is secured, without a system to provide the necessary context in a timely manner, validate execution results, and recover from failed tasks, it is difficult to operate reliably in a production environment.
Microsoft stated in a post published on June 2, 2026, that what enterprises need is not just strong models or computational resources but also systems to build, deploy, observe, and control agents. It presented context, governance, observability, and continuous improvement as key conditions for introducing agents into production environments.
OpenAI also explained in a post published on June 25, 2026, that the way agents are used is shifting from one-off queries to delegating long-term tasks. As task durations increase, the importance of validating intermediate results, restricting permissions, and handing over procedures to humans in case of failures also grows.
Security is another challenge that enterprise agents must address. As agents gain the ability to read and write internal documents and work systems, the risk of exposing sensitive information alongside the ability to utilize necessary context increases.
Microsoft Research's CI-Work study presented that the privacy violation rate of enterprise large language model (LLM) agents ranges from 15.8% to 50.9%, with information leakage rates reaching up to 26.7%. The researchers concluded that merely increasing model size or inference depth would not adequately resolve context-related information leakage issues.
In the industry, there are differing views on whether agents passing through screens and browsers should be seen as a transitional method to connect with existing software, or whether agents directly handling existing tools like humans will become mainstream. Both perspectives agree that access control, guardrails, and validation of execution results are necessary for actual implementation.
Efforts to establish operational structures for agents are also ongoing within companies. Coinbase's internal AI system has applied a cloud sandbox, coordination of subordinate agents, and an automatic task generation structure, which serves as a foundation for distinguishing execution scope and responsibilities as the model performs multi-step tasks.
In the virtual asset sector, the direct intersection lies in how to design identity, authority, and audit systems rather than the price benefits of AI agents. As joint research begins to verify the identity and permitted scope of autonomous trading AI agents, the importance of control systems increases as agents become more integrated into corporate tasks and digital asset infrastructure.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.





























