GeekWire Logo
Menu
  • Home
  • News
    • Amazon
    • Artificial Intelligence
    • Civic
    • Geek Life
    • Health/Life Sciences
    • Microsoft
    • GeekWire Podcast
    • Space
    • Startups sponsored by Google for Startups
    • Positive Charge Podcast sponsored by Amazon Sustainability
    • Sustainability sponsored by Amazon Sustainability
    • Tech Moves presented by GeekWork Recruiting
    • Agents of Transformation presented by Accenture
  • GeekWork
    • GeekWork Recruiting
    • Job Board
  • Events
    • Community Calendar
    • GeekWire Events
  • Lists
    • GeekWire 200
    • GeekWire Startup List presented by ALLtech
    • GeekWire Startup Resources
    • GeekWire Startup Spaces
    • Layoff Tracker
    • M&As and IPOs
    • Northwest Women VC & Angel Investor List
    • Recent Fundings
    • Seattle Engineering Outposts
    • Venture Capital Directory
  • Members
    • Health Benefits
    • Memberships
  • Studios
    • GeekWire Studios: Let Us Tell Your Story
    • AI Agents and Tools Podcast Series presented by AWS Marketplace
    • F5 30th Anniversary
    • AI Breakthrough Awards presented by UiPath
    • AI Innovators Spotlight Studio 2026 presented by AWS Marketplace
    • Nebius at NVIDIA GTC 2026
    • Acumatica Summit 2026
    • BMC Control Freaks Unite presented by BMC
    • AI Predictions Series 2026 presented by Unify Consulting
    • Guide to re:Invent 2025 sponsored by AWS
    • Remitly Reimagine 2025 sponsored by Remitly
    • Does Compute presented by Carnegie Mellon University
    • AWS Marketplace Seller Conference 2025 sponsored by AWS
    • ShopTalk 2025 powered by Amazon Ads
  • About
    • About GeekWire
    • Advertise
    • Contact Us
    • Email Newsletters
    • Reprints & Permissions
  • Podcast
  • LinkedIn
  • Newsletter
  • News Tips

What happens here matters everywhere.

  • Podcast
  • LinkedIn
  • Newsletter
  • News Tips
  • Amazon
  • Microsoft
  • Startups
  • AI
  • Science
  • Tech Moves
  • Sustainability
  • Civic
  • Geek Life
Sponsored Post

In AI we trust?

by Silvio Savarese on Jul 28, 2025 at 12:00 amAugust 6, 2025 at 9:37 pm

  • Facebook
  • X (Twitter)
  • LinkedIn
  • Email

A recent study by Stanford University’s Social and Language Technologies Lab (SALT) found that 45% of workers don’t trust the accuracy, capability, or reliability of AI systems. That trust gap reflects a deeper concern about how AI behaves when the stakes are high, especially in business-critical environments.

Hallucinations in AI may be acceptable when the stakes are low, like drafting a tweet or generating creative ideas, where errors are easily caught and carry little consequence. But in the enterprise, where AI agents are expected to support high-stakes decisions, power workflows, and engage directly with customers, the tolerance for error disappears. True enterprise-grade reliability demands more: consistency, predictability, and rigorous alignment with real-world context, because even small mistakes can have big consequences.

This challenge is referred to as “jagged intelligence,” where AI systems continue to shatter performance records on increasingly complex benchmarks, while sporadically struggling with simpler tasks that most humans find intuitive and can reliably solve. For example, a model might be able to defeat a chess grandmaster that is unable to complete a simple child’s puzzle. This mismatch between brilliance and brittleness underscores why enterprise AI demands more than general LLM intelligence alone; it requires contextual grounding, rigorous testing, and continuous fine-tuning.

That’s why at Salesforce, we believe the future of AI in business depends on achieving what we call Enterprise General Intelligence (EGI) – a new framework for enterprise-grade AI systems that are not only highly capable but also consistently reliable across complex, real-world scenarios. In an EGI environment, AI agents work alongside humans, integrated into enterprise systems and governed by strict rules that limit what actions they can take.

To achieve this, we’re implementing a clear, repeatable three-step framework – synthesize, measure, and train – and applying this to every enterprise-grade use case.

A Three-Step Framework for Building Trust

Building AI agents within the enterprise demands a disciplined process that grounds models in business-contextualized data, measures performance against real-world benchmarks, and continuously fine-tunes agents to maintain accuracy, consistency, and safety.

  • Synthesize: Building trustworthy agents starts with safe, realistic testing environments. That means using AI-generated synthetic data that closely resembles real inputs, applying the same business logic and objectives used in human workflows, and running agents in secure, isolated sandboxes. By simulating real-world conditions without exposing production systems or sensitive data, teams can generate high-fidelity feedback. This method is called “reinforcement learning” and is a critical foundation for developing enterprise-ready AI agents.
  • Measure: Reliable agents require clear, consistent benchmarks. Measuring performance isn’t just about tracking accuracy, it’s about defining what each specific use case requires. The level of precision needed varies: An agent offering product recommendations may tolerate a wider margin of error than one evaluating loan applications or diagnosing system failures. By establishing tailored benchmarks such as Salesforce’s initial LLM benchmark for CRM use cases, and acceptable performance thresholds, teams can evaluate agent output in context and iterate with purpose, ensuring the agent is fit for its intended role before it ever reaches production.
  • Train: Reliability isn’t achieved in a single pass — it’s the result of continuous refinement. Agents must be trained, tested, and retrained in a constant feedback loop. That means generating fresh data, running real-world scenarios, measuring outcomes, and using those insights to improve performance. Because agent behavior can vary across runs, this iterative process is essential for building stability over time. Only through repeated training and tuning can agents reach the level of consistency and accuracy required for enterprise use.

Turning AI Agents Into Reliable Enterprise Partners

Building AI agents for the enterprise is much more than simply deploying an LLM for business-critical tasks. Salesforce AI Research’s latest research shows that generic LLM agents successfully complete only 58% of simple tasks and barely more than a third of more complex ones.

Truly effective EGI agents that are trustworthy in high-stakes business scenarios require far more than an off-the-shelf DIY LLM plug-in. They demand a rigorous, platform-driven approach that grounds models in business-specific context, enforces governance, and continuously measures and fine-tunes performance. The AI we deploy in Agentforce is built differently. Agentforce doesn’t run by simply plugging into an LLM. The agents are grounded in business-specific context through Data Cloud, made trustworthy by our enterprise-grade Trust Layer, and designed for reliability through continuous evaluation and optimization using the Testing Center. This platform-driven approach ensures that agents are not only intelligent, but consistently enterprise-ready.

As businesses evolve toward a future where specialized AI agents collaborate dynamically in teams, ‌complexity increases exponentially. That’s why leveraging frameworks that synthesize, evaluate, and train agents before deployment is critical. This new framework builds the trust needed to elevate AI from a promising technology into a reliable enterprise partner that drives meaningful business outcomes.

Silvio Savarese is executive vice president and chief scientist at Salesforce AI Research.
  • Facebook
  • X (Twitter)
  • LinkedIn
  • Email

Latest Stories

    Read All Stories

    GeekWire Newsletters

    Subscribe to GeekWire's free newsletters to catch every headline

    Most Popular on GeekWire

      A Word From Our Sponsors

      About

      • About GeekWire
      • Contact Us
      • Partner With Us
      • Become a GeekWire Member
      • Send Us a Tip
      • Join Our Startup List
      • Reprints and Permissions

      Follow

      • Facebook
      • X
      • LinkedIn
      • Instagram
      • RSS Feed
      • Podcast
      • YouTube
      • Bluesky

      GeekWire Newsletters

      Catch every headline in your inbox

      Read GeekWire

      • Apple News
      • Google News

      Legal

      • Privacy Policy
      • Terms of Use
      • Sponsored Content Policy
      Return to Top of Page
      © 2011-2026 GeekWire, LLC
      Do Not Sell or Share My Personal information
      Limit the Use Of My Sensitive Personal Information
      Consent Preferences

      Don't miss a headline!

      Sign up for GeekWire's newsletters today.

      Thanks for subscribing!
      Check your inbox to confirm your subscription.