Episode 145

Target: Why AI Projects Fail Between Demo and Production

Sowmya Podila
Sowmya Podila
Sr. Data Scientist for Generative AI

In this episode we talked about:

  • Why AI demos can succeed while the same systems fail in production
  • What evaluation, governance, observability, and security practices enterprise AI needs to scale
  • Why narrowly scoped domain agents often perform better than agents designed to do everything
  • How existing processes and “golden datasets” make strong foundations for enterprise AI use cases
  • Why lower-risk, repeatable tasks are often the best place to begin AI adoption
  • How to evaluate whether AI automation actually delivers enough ROI to justify its cost
  • Why retailers increasingly need to think about discoverability inside AI agents, not just traditional search

🎧 Listen now on Apple Podcasts, Spotify, or YouTube

Episode highlights:

00:51 – Sowmya’s Journey From Data Science To Enterprise GenAI

‍03:28 – What It Takes To Scale AI Beyond The Demo

‍07:19 – Why Well-Scoped AI Agents Win In Enterprise

‍10:16 – Building An AI Agent For Faster Root Cause Analysis

‍15:01 – Why Retailers Need To Prepare For Agentic Shopping

Sowmya's bottom line: A demo only has to work once, but production is judged by a single failure across a million runs. At Target, that means treating AI as an engineering practice, with evaluations, observability, guardrails, and clear ownership, and focusing on well-scoped domain agents where established processes and golden datasets already define what good looks like.

FAQ

Sowmya Podila is a Senior Data Scientist for Generative AI at Target, where she has spent the last two years on a centralized advanced AI team that sets enterprise-wide AI strategy and embeds with product and business teams to take GenAI use cases from POC and MVP into production. She started her career as an engineer at Tata Consultancy Services in India, moved to the U.S. for a master's in technology management, and did machine learning in HR analytics at Gartner before joining AWS, where she worked in natural language processing and machine learning across clients in many domains. Generative AI, she says, became a natural extension of that work once ChatGPT launched.
Because the bar is completely different. Sowmya's framing: a demo only needs one good run for everyone to think it works, but production is judged by a single failure across even a million points. She points to reports showing 70% to 80% of AI use cases stall in the POC phase, or get rolled out and then rolled back, because they don't meet expectations at production level or fall short of governance and security standards.
Treating AI as an engineering practice rather than a prototype. Sowmya lists the questions enterprises now have to answer: what evaluations are needed post-demo to prove the solution works consistently at scale; what infrastructure scaling it takes to run with low latency in front of a large user base; what observability and guardrails will catch new failure modes before an agent goes rogue; and what governance applies when something goes wrong, including who owns it and how the company complies with federal and state regulations. Enterprises are also defining an agent taxonomy (what counts as an AI system, an agent, a super agent, or a sub-agent). The companies doing this work are the ones seeing stability, efficiency gains, and ROI.
In well-scoped domain agents rather than one agent trying to do everything. Sowmya's examples: within inventory management, making sure specific items never go out of stock and running root cause analysis when they do; within creative and fashion design, understanding trends or generating ad campaigns with AI imagery. The strongest candidates are processes humans have already run for years, where golden datasets define the inputs, the outputs, and what good looks like. That lets teams layer AI on top and measure how closely the AI output aligns with the human output.
An AI assistant for site reliability engineers. Large organizations often have one core SRE team supporting many products, so engineers can't hold every playbook in their heads; they open a guide, check dashboards, and drill down step by step to find the root cause within their SLAs, often handing work off between onsite and offshore teams. Sowmya's team put those playbooks behind a bot integrated into the tools engineers already use, such as Slack. When a failure is triaged, the bot acts as first responder: it reviews dashboards, suggests likely causes and next steps, completes what it can, flags steps that need a human, and generates a summary of what the engineer has done for the next team's hand-off.
Start with low-risk problems. Without enterprise-scale bandwidth for evaluation and guardrails, Sowmya advises against putting an AI assistant in front of customers where it could damage reputation or trust, or letting AI commit code to a production database and take the whole shopping site down. Generating an ad campaign or creating product imagery with AI instead of a photoshoot is a safer place to begin. The output may not be the highest-converting version, but the downside is contained.
In two places. First, areas that depend heavily on subject matter expertise and human decision-making, where AI hasn't lived up to performance expectations. Second, use cases where the economics don't work: if a job costs $10 to do manually and the AI spends $15, there's no gain. Sowmya notes vendors priced models low during the experimentation phase to drive adoption, and now that adoption is established, prices and token costs are rising, while many solutions rely on excessive context and too many LLM calls. Use cases where cost doesn't add up to ROI won't take off.
Because product discovery is moving into AI assistants, even when the purchase doesn't happen there. Sowmya explains that Target, working with Google on the Universal Commerce Protocol (UCP), was among the first retailers to integrate with ChatGPT, Microsoft Copilot, and Google. The goal is discoverability, which she says is now being called agent engine optimization in place of search engine optimization. Her own example as a new mom: she wouldn't necessarily buy baby bottles inside ChatGPT, but that's where she researched whether glass or plastic was safer and which options made sense. Exposing your catalog to these agents builds brand presence at that stage, and retailers are now competing for those slots the way they compete for top Google results, with sponsored placements likely to follow.

Want to be featured on The Ecommerce Toolbox?

We feature senior ecommerce leaders with real things to say. If you've led something worth talking about, we want to hear from you.

Apply to be a guest →

Other episodes

See all episodes

Audit your site today

Understand the revenue impact of all errors on your site and how to swiftly reproduce and resolve.