• Research finds AI agents haven't quite mastered real-world browsi

    From TechnologyDaily@1337:1/100 to All on Thursday, September 03, 2026 14:45:24
    Research finds AI agents haven't quite mastered real-world browsing tasks despite claiming they can

    Date:
    Thu, 03 Sep 2026 13:35:00 +0000

    Description:
    Many AI agents lack sufficient safeguards and can't handle multiple tabs very well despite promises.

    FULL STORY ======================================================================Copy link Facebook X Whatsapp Reddit Pinterest Flipboard Threads Email Share this article 0 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter Not a single agent scored the full 20 out of 20 Claude for Chrome performed better than the ChatGPT Chrome Extension With performance varying by testing category, Decodo advises selecting an agent based on planned usage New Decodo research has criticized AI agents for still not being able to conduct real-world browsing tasks autonomously, including tasks like form filling, completing transactions, having cross-tab awareness and handling third-party integrations.

    In fact, the study analyzed 45 AI agents across 10 different capabilities and found that not a single one could achieve the maximum score of 20. The
    testing is also said to have exposed discrepancies between what vendors and
    AI developers say their agents are capable of, and what they can actually deliver on. Latest Videos From TechRadar Watch full video here: Agentic AI isn't at the level of autonomous browsing, yet Claude for Chrome was the highest-scoring agent, reaching 18 points two below the theoretical maximum. Its OpenAI counterpart, the ChatGPT Chrome Extension, fell short with a 14-point score.

    Decodo argues that agents most commonly fall behind on transactions the lowest-scoring category with an average of 0.43 out of 2. The study found that, while they can often reach the checkout stage, they can't actually complete a purchase on behalf of the user. You may like Claude Sonnet 5
    booked nothing, and still felt like an assistant ChatGPT's new Side Chat features enhance browsing in Chrome New study reveals how AI is reshaping how people shop online but they still don't really trust it enough

    Even if an agent is technically capable of making a purchase, the paper warns that it still might not have the best safeguards in place to keep sensitive information like credit card numbers safe.

    Safeguards is a major theme in Decodo's work, with the research also calling out multi-step workflows for lacking sufficient safeguards. More than half of the agents capable of multi-step workflows lacked any documented safeguards before conducting irreversible actions. Are you a pro? Subscribe to our newsletter Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed! Contact me
    with news and offers from other Future brands Receive email from us on behalf of our trusted partners or sponsors By submitting your information you agree to the Terms & Conditions and Privacy Policy and are aged 16 or over.

    While current agentic technology still doesn't deliver on its promises, one type may emerge as the winner. Decodo found several browser-native agents to score full marks for cross-tab awareness.

    "Match the tool to the job, not to the longest feature list," Product Marketing Team Lead Gabriele Vitke explained, urging AI agent users to consider their own needs instead of buying into marketing. Follow TechRadar
    on Google News and add us as a preferred source to get our expert news, reviews, and opinion in your feeds.



    ======================================================================
    Link to news story: https://www.techradar.com/pro/research-finds-ai-agents-havent-quite-mastered-r eal-world-browsing-tasks-despite-claiming-they-can


    --- Mystic BBS v1.12 A49 (Linux/64)
    * Origin: tqwNet Technology News (1337:1/100)