Expert Perspectives
Expert Perspectives
Episode 145


In this episode we talked about:
- Why AI demos can succeed while the same systems fail in production
- What evaluation, governance, observability, and security practices enterprise AI needs to scale
- Why narrowly scoped domain agents often perform better than agents designed to do everything
- How existing processes and “golden datasets” make strong foundations for enterprise AI use cases
- Why lower-risk, repeatable tasks are often the best place to begin AI adoption
- How to evaluate whether AI automation actually delivers enough ROI to justify its cost
- Why retailers increasingly need to think about discoverability inside AI agents, not just traditional search
🎧 Listen now on Apple Podcasts, Spotify, or YouTube
Episode highlights:
Sowmya's bottom line: A demo only has to work once, but production is judged by a single failure across a million runs. At Target, that means treating AI as an engineering practice, with evaluations, observability, guardrails, and clear ownership, and focusing on well-scoped domain agents where established processes and golden datasets already define what good looks like.
Sowmya Podila & Kailin Noivo — Transcript
The Ecommerce Toolbox: AI in Retail • Human-Reviewed Transcript
[00:00:00] Sowmya Podila: Hey, demos actually need one good run, and you think everything is working. But when you wanna take it into production, you actually would judge the quality of production by one failure across even a million points.
[00:00:11] Kailin Noivo: Welcome to another episode of The Ecommerce Toolbox: AI and Retail. Joining us today, we have a very special guest, Sowmya Podila, and she is the senior data scientist for generative AI over at Target. So, welcome, Sowmya.
[00:00:26] Sowmya Podila: Yeah. I'm glad to be here, Kailin. I have seen a lot of interesting conversations on the ecommerce podcast.
[00:00:31] Kailin Noivo: Yeah. No. Exactly. We just rebranded it, actually. We're now just a full-blown AI in retail podcast. But we always like to start off, Sowmya, by asking a bit about your career journey. You've had a really cool career. You've worked at places like Gartner and AWS before Target. So, maybe take us through your career journey and how you ended up in your current role.
[00:00:50] Sowmya Podila: Yeah. Definitely. I started my career actually at Tata Consultancy Services, TCS, back in India, and I was an engineer at that time. We moved to the U.S. here to do my master's in technology management. That's where I got introduced to data science as a sweet spot between, you know, tech and business. And since then, I have been in love with data science and started actually working in machine learning. And very interestingly, I worked in HR analytics and did machine learning in HR at Gartner, which was a very unique niche space at the time. And later, I wanted to diversify my skill set in terms of both the tech and also working in different domains, and that brought me to AWS working on the cloud stack and working across a lot of cloud services. Getting that depth into the technology stack and working with various AWS clients across domains helped me gain understanding into how each domain could be different from each other and how we can still apply tech across domains. So, that was a very strong experience. And I was working in natural language processing and machine learning, and then when ChatGPT launched and all those kinds of generative AI technologies became mainstream, that became a natural extension of my work. On those lines, Target came as a new opportunity for me where they were building a centralized advanced AI team where you gathered a bunch of AI experts and you are doing enterprise-wide AI strategies and embedding into various product and business teams inside Target, helping them enable with GenAI upskilling and use cases and doing, like, a test run, build of POC style and MVP, and then taking use cases that kind of live up to all the expectations and, you know, move into production. So, that was the approach we were doing at Target. And that kind of plan a couple of years ago on how they want to tackle this solutioning in AI really impressed me, and that's why I'm here at Target doing it for the last two years. Yeah. So, from computer science undergrad, AI/ML, now on to GenAI has been my career arc. Yeah.
[00:02:45] Kailin Noivo: That's really cool. And you were doing AI before AI was really cool. Now everybody wants to do AI. Tell us a bit about what you're seeing actually work in the enterprise. So, we're seeing a lot of hype, obviously, in the space. And demos and AI, they all kinda look the same. You click a button, and then all these keyboards start to stroke themselves, code being written, changes are being made, everyone's making money. How are you seeing that play out in the field in enterprise? And how is that different from what you've seen in maybe some smaller retailers?
[00:03:18] Sowmya Podila: Yeah. Definitely. No. That's kind of my perception of how a lot of AI hype is generated as well in the industry as a whole. I kind of had a line saying that, "Hey. Demos actually need one good run, and you kind of think everything is working. But when you wanna take it into production, you actually would judge the quality of production by one failure across even a million points." Right? So, you know, that's how the difference is, you know, when you want to take a demo into production. So, that's where a lot of use cases are getting stalled. So, if you have seen a lot of reports that are released, I would say they talk about 70% to 80% of the use cases getting stalled in POC phase or even rolled out and then rolled back because they are not meeting expectations at a production level, or are not meeting the governance standards or security standards as well. So, yeah, that's been the reality for sure. And now there has been a lot of emphasis or focus on understanding this as an engineering infrastructure or engineering practice, you know, that could work at an enterprise scale. So, what are the evaluations that you would need to place post-demo to ensure that this would work consistently at scale and in production? So, what is the infrastructure scaling needed for it to run performantly with low latency when you put it in front of a big set of users? And what is the observability and guardrails needed to ensure that we monitor the solution in production and catch it before the agent goes rogue with any new kind of failure modes? And also, what are the governance pieces needed when something is not right? Who is the responsible party or owner for it? You know? And other kinds of responsible AI concerns, ethical concerns. How do we ensure we comply with the federal and state regulations as well? So, all these aspects are now thought through. So, there are a lot of buzzwords as well. So, one of the other aspects is coming up with the agent taxonomy, defining what is AI, what is an AI agent, what is an AI system, what is a super agent, what are sub-agents, you know, all these kinds of taxonomies are also getting now defined at an enterprise level. So, all of these enterprise-level, engineering-perspective things that are needed for an AI demo to be successful in production are what enterprises are now focusing on, and companies focusing on that are seeing more success and stability, maybe seeing ROI at the end of the day or efficiency gains at least partially. So, these practices are what are making the use cases successful.
[00:05:57] Kailin Noivo: Yeah. It makes sense. And, like, just diving into the categories. So, at least, what we're seeing, we're building a lot of ecommerce agents, and they're being used across SMB over to enterprise. But to your point, enterprise is a lot more governance. So, in some of our smaller clients, like a traditional Shopify store, they might not even have a development team. Maybe they're using someone offshore. It's not uncommon that they run our agents. The agents come up with a PR in GitHub. They click the button. It releases it. That's it. Like, they're just making changes. Right? And then, obviously, as you climb up the ladder from mid-market into enterprise, we're seeing people start to have CI/CD pipelines and a staging environment. Some of these smaller guys are just launching it into the wild, making changes. And as they go up to enterprise, more and more humans get brought into the loop. Right? And it's not uncommon that some of our end-to-end workflows in enterprise go to a staging website and have multiple human interventions. So, I guess my question for you is, like, in which categories are you seeing in enterprise the most adoption? Right? Is it, for example, regression testing? Is it, like, bug fixing? Is it compliance? Like, where are you seeing? Is it content generation? Is it customer support? Where are you seeing the most use cases, out of curiosity, in the enterprise specifically?
[00:06:58] Sowmya Podila: Yeah. So, I'm kind of seeing more success with well-scoped projects such as domain agents. For example, you take one department, such as inventory management, and one specific use case even inside that inventory management, such as you want to understand the items that you would never wanna run out of stock. So, how do we ensure that those items are always in stock? And if they went out of stock, how can we do root cause analysis with AI agents, for example? That is one specific use case. Another use case could be, like, you know, within fashion designing or creative designing, how can we ensure efficient understanding of trends or efficient generation of an ad campaign through, you know, AI-generated imagery. So, a very well-scoped specific use case in any of the domains is where I kind of see success, rather than trying to do it all with one agent or do too many things at once. So, those are the kind of domain agent use cases where I'm seeing a lot of success. And inside the domain, like I said, to expand on the scope, redundant use cases where you already have established processes and what good looks like, or golden datasets. Hey, humans have done this for 5 years. This is what they have done across these 5 years. We already have the dataset. This is the input. This is the output. This is what good looks like. That is all defined. Now, if we already have the established process, how can we now add an AI layer on top of it to automate some of this? And now we have all the golden datasets to run it through this AI pipeline to ensure that the human output versus AI output — how synergistic are these? How well aligned are these? Right? So wherever you have these well-scoped existing datasets or processes or methodology to evaluate what good looks like, those are the use cases where we see a lot of success. So, in terms of smaller retailers or other organizations that do not have that big of a bandwidth to kind of build at an enterprise scale, what I would suggest is to pick the lower-risk things. For example, you do not want to put an AI assistant in front of your customers that would damage your reputation or trust, or you do not wanna commit code on your production-grade database and bring your whole shopping website down so that you would never be able to recover. You do not want to take that big of a risk, right? But if you're generating an ad campaign or, instead of taking pictures of your products, you're designing them with AI. So, it might not be the product with the most conversions, but you are not running into bigger risks for yourself, right? So, I would suggest choosing low-risk problems for those kinds of organizations as well.
[00:09:27] Midroll: If you're listening to The Ecommerce Toolbox, you're entitled to a podcast-exclusive website audit. Go to noibu.com/podcast-audit for a free scan that uncovers the hidden friction blocking your conversions and shows you where you're leaking revenue.
[00:09:41] Kailin Noivo: So, I was reading ahead of time that you had actually built an SRE agent. So, maybe talk to me a bit about that, and then how that drove productivity lift?
[00:09:50] Sowmya Podila: Yeah. So, I think every other company, especially one that has a bigger public-facing infrastructure, would have a site reliability engineering team. Right? So, there would be a lot of failures in staging or production that, you know, these engineers need to address, and there is a lot of material that helps them track down, when they see an error or failure, where it could be arising from. So, there's a lot of work that goes into root cause analysis within the SLAs to figure out where the production failure is coming from and to be able to fix it. A lot of times, for organizations that are big, there are a lot of products, and one core team is supporting a lot of products at the same time, where they will not have all the needed insights or guides in their head. So, they have to open a guide or a playbook that they walk through step by step to figure out the root cause. First, they go to a dashboard, try to see where the failures have happened, and, you know, drill down on one failure point and then keep drilling down till they figure out the answer. Right? And a lot of people actually work around the clock. There is an onsite team and an offshore team, and a lot of hand-offs happening as well. So, that's where we kind of thought, you know, why don't we take all these guides and playbooks that exist across the countries for a lot of products together and put them behind an AI assistant and integrate it with where they're already working. So, they do this root cause analysis as a team together, probably on Slack. And when the clock ticks down, they kind of hand it off to the offshore team and sign off. So, we built an AI assistant or a bot and integrated it into these platforms. So, whenever there's a failure that's been triaged, the bot acts as the first point of response, guiding the SRE engineers to help them track down the production failure or the root cause faster. So, it takes as input all these playbooks and then helps them: "Hey, I already saw some of these dashboards. I kind of see that this might be the issue. These are the next few steps. Would you wanna start with these? Or I already did two of them, but two of them need human intervention. Would you like to perform these two steps?" So, it's that kind of assistant for the engineers to be able to, you know, effectively trace down production failures and also hand off across cross-country teams. Basically, AI will generate the summary of what an engineer has done and hand it off, you know, to the next team person. So, that's the kind of intent behind these kinds of initiatives.
[00:12:01] Kailin Noivo: No. That's honestly really powerful. And I think you're right in saying, like, a lot of enterprises have just too much risk with their reputation to effectively just have, like, some customer-facing agent go to town. Where do you think agents are gonna get the least amount of adoption in the next year in the enterprise?
[00:12:20] Sowmya Podila: That's a very good question. So, I would say that areas where a lot of subject matter expertise and human decision-making are needed is one area where we have tried it, and it's not living up to the performance expectations. And the other area is, yeah, AI is able to do it, but for a job that costs $10 to do manually, the AI is spending $15, actually. There is not actually a revenue gain from implementing AI. So, initially, when we were all in the experimentation phase, even the vendors were pricing these models so low to kind of boost the adoption. Now that there is a good amount of adoption, the prices or the token costs, or the economics, are not adding up. The prices are going up, and the solutions are kind of built around using excessive context, making too many LLM calls to arrive at an answer, and the cost of that is increasing. So, wherever there are use cases where, you know, the cost is not adding up to the ROI, those might not be the ones that would take off for sure.
[00:13:20] Kailin Noivo: No. I'm fully aligned. I think, honestly, where it's getting the most traction is in mundane tasks that are easily repeatable that you would generally hire a junior team to do. Like you said, first-line bug triage support or investigating a latency spike or something to that tune. Talk to me a bit about how you think it's gonna be important for brands like Target or other brands to open up their catalog to third-party AI agents to scrape. It's really interesting because I think, like, I will never buy a luxury handbag for my wife through AI, but, like, I broke a light bulb in my house, and it is a very strangely shaped light bulb. And I took a picture of it, and it was like, I need to buy one, like, figure out how I can get one in two days. And it was like, tu-tu-tu. So, I feel like there's gonna be some use cases for sure for AI. So, maybe talk to me like, sorry, like, agentic shopping. So, maybe talk to me a bit about how important it is for the catalog, and are there some security risks in that?
[00:14:24] Sowmya Podila: No. Definitely. I think that's how I see it too. Even if we do not buy these kinds of, you know, fancy purchases, if you want to discover what is a fancier purchase or what would make a good handbag or what color option, what is trending out there, especially when you do not have knowledge for the discovery, you might still go to these platforms to get a sense. Right? So, you know, I think it's kind of — so, for context, Target has also developed, in partnership with Google, the Universal Commerce Protocol (UCP), that has been the backbone of how these kinds of retailers are integrating with these apps. And Target has been one of the first companies to race and integrate with ChatGPT and Microsoft Copilot, along with Google as well, developing this protocol and pioneering this. I think it is coming from a sense of being discoverable, as they're calling it. Instead of search engine optimization, now they're calling it agent engine optimization. So, Google was a search where you click. You might not complete the purchase itself. You might come to the app and search more, but, again, you want to be discoverable on Google. Right? Similarly, now people are not spending as much time purely on Google as before and browsing through catalogs, but asking these AI agents to kind of figure out what they should be buying. For example, I'm a new mom, and walking into motherhood, I had no clue what to buy for my baby. Would I complete the purchase on ChatGPT? No. But, hey, what kind of bottles are safe? How do I decide between plastic versus glass feeding bottles for the baby? And all these questions and all of the discovery browsing and informed perspective and maybe just seeing through some options to kind of, you know, paint the picture in my head on which of these make sense for me versus not. That discovery has been happening on these platforms. So, even if conversions are not happening, definitely, building a brand image and kind of, you know, giving more exposure to your catalog through these agents is definitely the need of the hour for all the retailers, and everyone is competing for that space now, the way they also compete for the top slots on Google search. Now, similarly, all the retailers are competing for that. And now, with the concept of sponsored ads as well on these platforms, I am not sure how it's all gonna shape up, but eventually, retailers might catch up and do that as well, you know.
[00:16:31] Kailin Noivo: Funny you say that. One, congratulations. My wife and I just had a baby as well, and very similar. My wife every single day is asking a different question. Is this normal? Is this not normal? Like, should I be concerned? What do I do? Are these chunks too big? Do I need to blend it more? You know how it is. So, yeah, very cool. I appreciate the time. This was a phenomenal conversation. Obviously, huge congratulations on all your success at Target and your previous businesses, and I learned a lot about enterprise AI deployment. So, I really, really appreciate you taking the time today.
[00:17:06] Sowmya Podila: Yeah. Likewise. Thank you so much. It's a pleasure to join the conversation with you and share some of what I'm seeing here with the broader audience.
[00:17:13] Kailin Noivo: Awesome. Thank you so much.
[00:17:16] Sowmya Podila: Thank you. Bye bye.
[00:17:16] Outro: The Ecommerce Toolbox: AI and Retail is brought to you by Noibu. To find out more about Noibu and how we unify error monitoring, site performance, and experience analytics to uncover growth opportunities and skyrocket your revenue, visit www.noibu.com. That's n-o-i-b-u.com. And then make sure to search for The Ecommerce Toolbox: AI and Retail on Apple Podcasts, Spotify, or anywhere else podcasts are found, and click subscribe so you don't miss out on any future episodes. On behalf of the team here at Noibu, thanks for listening.
FAQ
Other episodes
Audit your site today
Understand the revenue impact of all errors on your site and how to swiftly reproduce and resolve.




.png)
.png)
.png)
.png)
.png)

.png)
.png)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)

.avif)
.avif)
.avif)
.avif)
.avif)

.avif)

.avif)

.avif)
.avif)
.png)
.avif)
.avif)
.avif)

.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)

.avif)
.avif)
.avif)

.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)

.avif)
.avif)
.avif)

.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)

.avif)
.avif)
.avif)
.avif)
.avif)
.avif)
.avif)

.avif)
.avif)