Tata digital

AI Search

Product designer

2023

0 → 1

10 min read

Designed and deployed search co-pilot for Tata Neu, an AI-powered conversational assistant built in partnership with Slang Labs to bridge the gap between unstructured human intent and structured e-commerce catalogue.

Designed and deployed search co-pilot for Tata Neu, an AI-powered conversational assistant built in partnership with Slang Labs to bridge the gap between unstructured human intent and structured e-commerce catalogue.

G-Eval score

4.58

(from 4.42)

Null results

4%

(from 9%)

Co-pilot CTR

41.4%

(Search CTR was at 10%)

executive summary

my role

product designer

project lead

akshita rastogi

in collaboration with

SlangLabs

categories

electronics, fashion, grocery, health (EFGH)

This case study outlines the end-to-end product design and strategic execution of Search Co-Pilot for the Tata Neu super app. Developed in collaboration with Slang Labs using their CONVA.ai platform, this in-app digital assistant was designed to bridge the gap between unstructured human intent and structured e-commerce search engines. Operating across four primary categories - Electronics, Fashion, Grocery, and Health (EFGH), the Search Co-Pilot enables conversational, multimodal (voice and text), and multi-category product discovery.

This case study outlines the end-to-end product design and strategic execution of Search Co-Pilot for the Tata Neu super app. Developed in collaboration with Slang Labs using their CONVA.ai platform, this in-app digital assistant was designed to bridge the gap between unstructured human intent and structured e-commerce search engines. Operating across four primary categories - Electronics, Fashion, Grocery, and Health (EFGH), the Search Co-Pilot enables conversational, multimodal (voice and text), and multi-category product discovery.

AI Search journey

AI Search journey

problem space

ambiguity in standard search

Standard search engines demand precision. Users must translate their actual needs into rigid, database-friendly keywords.

rigid query parsing

Traditional engines fail to handle user misspellings, semantic synonyms, or natural language requests cleanly

Traditional engines fail to handle user misspellings, semantic synonyms, or natural language requests cleanly

UX friction

When a user searches for "what AC should I buy for my 10x10 room?" traditional parsing fails, leading to null results or incorrect category routing.

When a user searches for "what AC should I buy for my 10x10 room?" traditional parsing fails, leading to null results or incorrect category routing.

wrong category classification

Standard search engines frequently misclassify complex or vague queries, routing users to incorrect landing pages

Standard search engines frequently misclassify complex or vague queries, routing users to incorrect landing pages

rigid query

rigid query parsing resulting in null search

rigid query

rigid query parsing resulting in null search

auto suggestion

Manual category classification

auto suggestion

Manual category classification

to bridge the gap

to bridge the gap

aligning conversational intent with structured catalog retrieval to close the gap between what users want and what the system understands

solution space

strategic opportunity

create a contextual, intelligence-led wrapper over the existing Tata Neu search architecture

opportunity

Implementing a multimodal intelligence layer capable of understanding natural language, auto-classifying intent, and explicitly disambiguating complex queries.

Implementing a multimodal intelligence layer capable of understanding natural language, auto-classifying intent, and explicitly disambiguating complex queries.

increased Search Adoption

Increasing search engine result CTR & unlocking need-based ("What should I buy for a trek to Sikkim?") and assisted search behaviors (parsing long list items like "milk, bread, butter, jam")

Increasing search engine result CTR & unlocking need-based ("What should I buy for a trek to Sikkim?") and assisted search behaviors (parsing long list items like "milk, bread, butter, jam")

reduction in null results

transitioning query translation away from rigid keyword matching toward Generative AI context-mapping

transitioning query translation away from rigid keyword matching toward Generative AI context-mapping

framework & timeline

search framework

hybrid based on search intent

timeline

4 months (starting Oct 2023)

technical architecture

To balance cost efficiency, LLM response latency, and system reliability, we established a hybrid architectural framework. Search Co-Pilot acts as an intelligence wrapper over the native Tata Neu Global Search.

structured queries

Handled directly by the native search engine to minimize computation overhead across the app's daily queries

Handled directly by the native search engine to minimize computation overhead across the app's daily queries

Unstructured Intent (The Co-Pilot Layer)

Unstructured input is formatted and processed by the Co-Pilot to identify the intent and category. It generates structured search parameters (e.g., Category: Fashion, Keywords: Watches, Wallets) and hands them off seamlessly to the underlying partner brand APIs (BigBasket, 1mg, Croma, Tata CliQ).

Unstructured input is formatted and processed by the Co-Pilot to identify the intent and category. It generates structured search parameters (e.g., Category: Fashion, Keywords: Watches, Wallets) and hands them off seamlessly to the underlying partner brand APIs (BigBasket, 1mg, Croma, Tata CliQ).

hybrid search architecture

timeline

Leading the design strategy meant defining a phased, scalable rollout that aligned technical readiness with UX maturity. The timeline was structured to validate assumptions early.

month 1 (foundation)

Deployed an off-the-shelf implementation without Tata Digital (TD) dataset training to test baseline UX interventions and UI branding

Deployed an off-the-shelf implementation without Tata Digital (TD) dataset training to test baseline UX interventions and UI branding

month 2 (cross-functional testing)

Transitioned to a CFT-ready (internal) build, optimizing the model with TD catalog data and filtering logic to test interaction quality and navigation.

Transitioned to a CFT-ready (internal) build, optimizing the model with TD catalog data and filtering logic to test interaction quality and navigation.

month 3 & 4 (Scale)

Propagated the app for external pilot testing with the Closed User Group (CUG)

Propagated the app for external pilot testing with the Closed User Group (CUG)

approach, execution

visual interface

Visually separate from standard search

scope

deep-link redirection, recipe search, Customer support redirection

Interaction Architecture & System Design

To build trust and clearly delineate the generative experience from standard keyword search, we created two distinct search experiences. The architecture required establishing clear UI states that communicated system status during high-latency AI operations

Design pillars

Design pillars guided the visual development of the Co-Pilot interface to ensure it was highly intuitive

Design pillars guided the visual development of the Co-Pilot interface to ensure it was highly intuitive

Clear AI identity

Visual language should signal generative intelligence

Visual language should signal generative intelligence

Multimodal fluency

Text, chips, voice interactions co-exist without clutter

Text, chips, voice interactions co-exist without clutter

visual separation

Clear visual separation differentiating AI from native global search

Clear visual separation differentiating AI from native global search

Utility over conversation

Strong search affordance reinforcing utility over conversational novelty

Strong search affordance reinforcing utility over conversational novelty

Invocation

Invocation

Process

Process

Response

Response

SERP

SERP

Key executions

Invocation entry point

The invocation icon was systematically placed outside but immediately adjacent to the primary search bar. This prevents mental models from confusing it with standard keyword search while maintaining prominent accessibility

The invocation icon was systematically placed outside but immediately adjacent to the primary search bar. This prevents mental models from confusing it with standard keyword search while maintaining prominent accessibility

Invocation

Invocation

typing search

typing search

voice search

voice search

Exploration

Exploration

Exploration - single search

Exploration - single search

Intent-to-navigation loop

For transactional or account-based queries (e.g., "I want a Tata Neu Credit card"), the interface avoids sending the user to a typical Search Engine Results Page (SERP). Instead, the system alerts the user and directly deep-links them to the corresponding interactive app core page (e.g., the NeuCard application section).

For transactional or account-based queries (e.g., "I want a Tata Neu Credit card"), the interface avoids sending the user to a typical Search Engine Results Page (SERP). Instead, the system alerts the user and directly deep-links them to the corresponding interactive app core page (e.g., the NeuCard application section).

Query

Query

Process comm

Process comm

Redirected page

Redirected page

Graceful Degradation

For out of scope intent, when users input non-supported or completely abstract queries (e.g., general knowledge questions like "How many gears does a car have?"), the UI flags the constraint. It presents helpful alternatives or passes the user to Niya, the core digital support assistant, rather than hitting a dead-end.

For out of scope intent, when users input non-supported or completely abstract queries (e.g., general knowledge questions like "How many gears does a car have?"), the UI flags the constraint. It presents helpful alternatives or passes the user to Niya, the core digital support assistant, rather than hitting a dead-end.

OOS Query

OOS Query

Response

Response

Support chat

Support chat

FAQ

FAQ

managing interactions

core philosophy

Yes, And

The "Yes, And" Approach to Conversational UI

A core philosophy in managing user interactions was ensuring no query resulted in a frustrating dead end, much like the "yes, and" principle in communication.

Unsupported Queries

If a user asked a general knowledge question out of scope, the system gracefully degraded. Instead of a blank screen, it displayed, "Looks like this is beyond our capability to answer," and immediately provided actionable pathways to FAQs or the support Chatbot

If a user asked a general knowledge question out of scope, the system gracefully degraded. Instead of a blank screen, it displayed, "Looks like this is beyond our capability to answer," and immediately provided actionable pathways to FAQs or the support Chatbot

Direct Transactional Routing

For internal Tata Neu features and Financial Services (FS), we bypassed the traditional Search Engine Results Page (SERP) entirely. If a user requested a credit card, the UI acknowledged with "Taking you to Tata Neu Credit Card," hid unnecessary suggestion chips, and automatically deep-linked them to the application flow

For internal Tata Neu features and Financial Services (FS), we bypassed the traditional Search Engine Results Page (SERP) entirely. If a user requested a credit card, the UI acknowledged with "Taking you to Tata Neu Credit Card," hid unnecessary suggestion chips, and automatically deep-linked them to the application flow

Real-world challenges

phase

alpha testing

System refinement

During the Cross-Functional Team (CFT) alpha testing phases, several critical user experience and edge-case classification issues emerged, highlighting the complexity of refining a generative interface

During the Cross-Functional Team (CFT) alpha testing phases, several critical user experience and edge-case classification issues emerged, highlighting the complexity of refining a generative interface

Intent mismatch & contextual blunders

Problem

Searching for "immunity boosting food" inside the grocery vertical generated an irrelevant top chip leading to dog food instead of human nutrition. Similarly, a search query for "colic friendly baby food" erroneously served a product chip for floor disinfectant cleaner

Searching for "immunity boosting food" inside the grocery vertical generated an irrelevant top chip leading to dog food instead of human nutrition. Similarly, a search query for "colic friendly baby food" erroneously served a product chip for floor disinfectant cleaner

Solution

Implemented a "common sense" semantic filter layer. Human-first consumption became the default programmatic parameter unless terms explicitly specified pet or home-care verticals

Implemented a "common sense" semantic filter layer. Human-first consumption became the default programmatic parameter unless terms explicitly specified pet or home-care verticals

Over-technical and complex chip output

Problem

When searching for "OTC medicines for skin rashes," the GenAI model generated highly clinical, technical terms like "hydrocortisone cream" as the selectable button chip. This forced users to double-back to Google to interpret the results

When searching for "OTC medicines for skin rashes," the GenAI model generated highly clinical, technical terms like "hydrocortisone cream" as the selectable button chip. This forced users to double-back to Google to interpret the results

Solution

The generation prompt rules were updated to emphasize consumer-friendly vocabulary, prioritizing actions and clear category descriptions rather than dense medical or technical terminology.

The generation prompt rules were updated to emphasize consumer-friendly vocabulary, prioritizing actions and clear category descriptions rather than dense medical or technical terminology.

Acoustic & linguistic vulnerabilities in voice search

Problem

Voice searches for "stainless steel sipper for toddlers" were frequently transcribed as "stainless steel slippers for toddlers", leading to illogical product associations. Similarly, "Eye drops for kids" translated directly into literal text as "I drops for kids", yielding zero-state null screens.

Voice searches for "stainless steel sipper for toddlers" were frequently transcribed as "stainless steel slippers for toddlers", leading to illogical product associations. Similarly, "Eye drops for kids" translated directly into literal text as "I drops for kids", yielding zero-state null screens.

Solution

The semantic engine was enhanced to cross-reference ambiguous phonetic outputs with current platform inventory data. "Eye drops" or "sippers" automatically override literal spellings based on active catalog relevance.

The semantic engine was enhanced to cross-reference ambiguous phonetic outputs with current platform inventory data. "Eye drops" or "sippers" automatically override literal spellings based on active catalog relevance.

G-Eval quality score

Quality improvement

4.42 → 4.58

Quality score gate

4.5

Understanding the G-Eval Score

The G-Eval score is a highly structured, human-evaluation-aligned framework that treats generative AI outputs as a pipeline rather than a single black box. Instead of simply guessing whether an AI response is "good," a larger, highly capable evaluation model (such as a scaled GPT model) splits every single Co-Pilot output into 7 distinct generation milestones:

The G-Eval score is a highly structured, human-evaluation-aligned framework that treats generative AI outputs as a pipeline rather than a single black box. Instead of simply guessing whether an AI response is "good," a larger, highly capable evaluation model (such as a scaled GPT model) splits every single Co-Pilot output into 7 distinct generation milestones:

7 milestones of evaluation model

What an Improvement from 4.42 to 4.58 Means to the Business

The business established a strict quality gate: the Co-Pilot could not be shown to real users unless it maintained an average score above 4.5. Moving to 4.58 unlocked the approval to launch a live Proof of Concept (PoC) with 5,000 active users.

A score of 4.42 meant that a noticeable percentage of user queries were still hitting intermediate model errors or awkward outputs. At 4.58, the engine demonstrated incredibly high precision across foundational data points.

The business established a strict quality gate: the Co-Pilot could not be shown to real users unless it maintained an average score above 4.5. Moving to 4.58 unlocked the approval to launch a live Proof of Concept (PoC) with 5,000 active users.

A score of 4.42 meant that a noticeable percentage of user queries were still hitting intermediate model errors or awkward outputs. At 4.58, the engine demonstrated incredibly high precision across foundational data points.

Designing for AI: How system level interventions drove the G-Eval Score

To elevate the G-Eval score from 4.42 to 4.58, we had to design the conversational constraints. Every LLM output was evaluated across 7 steps (Continuation, Correction, Intent, Category, Search Term, Suggestion Chips, Message). We used qualitative user friction to redesign how the AI formulated those specific steps.

To elevate the G-Eval score from 4.42 to 4.58, we had to design the conversational constraints. Every LLM output was evaluated across 7 steps (Continuation, Correction, Intent, Category, Search Term, Suggestion Chips, Message). We used qualitative user friction to redesign how the AI formulated those specific steps.

UI to API constraint mapping (optimizing chip length)

UX Friction

In early CFT versions, the AI generated highly descriptive, long-tail search chips (sometimes containing hyphens or complex need-based phrases). While these were linguistically impressive, they failed when passed to the underlying partner brand APIs, triggering null search results.

In early CFT versions, the AI generated highly descriptive, long-tail search chips (sometimes containing hyphens or complex need-based phrases). While these were linguistically impressive, they failed when passed to the underlying partner brand APIs, triggering null search results.

Design intervention

We redesigned the output constraints for Step 6 (Suggestion Chips). We forced the LLM to shorten the chips and strip incompatible characters, ensuring the generated UI elements seamlessly triggered the native Search Engine Results Page (SERP) without breaking the backend.

We redesigned the output constraints for Step 6 (Suggestion Chips). We forced the LLM to shorten the chips and strip incompatible characters, ensuring the generated UI elements seamlessly triggered the native Search Engine Results Page (SERP) without breaking the backend.

Cognitive load reduction (De-jargoning the output)

UX Friction

The LLM was occasionally too smart for its own good. For example, a user searching for "OTC medicines for skin rashes" on Tata 1mg received a highly clinical suggestion chip for "hydrocortisone cream". This forced users to leave the app and Google the medical term just to understand the recommendation.

The LLM was occasionally too smart for its own good. For example, a user searching for "OTC medicines for skin rashes" on Tata 1mg received a highly clinical suggestion chip for "hydrocortisone cream". This forced users to leave the app and Google the medical term just to understand the recommendation.

Design intervention

We established a strict UX principle: GenAI is supposed to disambiguate, not complicate or confuse. We redesigned the system prompt to prioritize consumer-friendly vocabulary over clinical terminology. This directly improved the G-Eval scores for Step 7 (Message) and Step 6 (Suggestion Chips) by making the outputs definitively more "Helpful".

We established a strict UX principle: GenAI is supposed to disambiguate, not complicate or confuse. We redesigned the system prompt to prioritize consumer-friendly vocabulary over clinical terminology. This directly improved the G-Eval scores for Step 7 (Message) and Step 6 (Suggestion Chips) by making the outputs definitively more "Helpful".

Desigining semantic guardrails (the "common sense" baseline)

UX Friction

The AI initially lacked real-world pragmatism. A search for "immunity boosting food" surfaced dog food as the top chip, and a search for "colic friendly baby food" bizarrely suggested floor disinfectant cleaner.

The AI initially lacked real-world pragmatism. A search for "immunity boosting food" surfaced dog food as the top chip, and a search for "colic friendly baby food" bizarrely suggested floor disinfectant cleaner.

Design intervention

We designed a contextual defaulting logic. For example, human food became the absolute baseline assumption for grocery searches unless a user explicitly stated otherwise. Designing these rigid contextual boundaries vastly improved the accuracy of Step 3 (Intent Identification) and Step 4 (Category).

We designed a contextual defaulting logic. For example, human food became the absolute baseline assumption for grocery searches unless a user explicitly stated otherwise. Designing these rigid contextual boundaries vastly improved the accuracy of Step 3 (Intent Identification) and Step 4 (Category).

Multimodal error forgiveness (Acoustic parsing)

UX Friction

Voice inputs frequently resulted in literal, phonetic mistranslations. "Eye drops for kids" was parsed as "I drops for kids," and "stainless steel sipper" became "stainless steel slippers". Because the AI took the text literally, it returned zero results.

Voice inputs frequently resulted in literal, phonetic mistranslations. "Eye drops for kids" was parsed as "I drops for kids," and "stainless steel sipper" became "stainless steel slippers". Because the AI took the text literally, it returned zero results.

Design intervention

We designed a semantic cross-referencing loop for Step 2 (Is Corrected). Rather than blindly trusting the raw voice-to-text output, the system evaluates phonetically ambiguous phrases against the logical Tata Neu catalog (e.g., recognizing that "I drops" in a pharmacy context is an error).

We designed a semantic cross-referencing loop for Step 2 (Is Corrected). Rather than blindly trusting the raw voice-to-text output, the system evaluates phonetically ambiguous phrases against the logical Tata Neu catalog (e.g., recognizing that "I drops" in a pharmacy context is an error).

measurable impact

experiment timeframe

60 days

scale

5000 users

Impact of testing

The experiment ran for over 60 days, processing more than 5,000 user queries across app versions 5.0.2 and 5.1.0

The impact by numbers

55 issues resolution

Approximately 55 issues related to intent understanding, performance, and stability were identified and fixed during this period

Approximately 55 issues related to intent understanding, performance, and stability were identified and fixed during this period

G-eval quality score from 4.42 to 4.58

The LLM's performance score measured on a scale of 1 to 5 for helpfulness, factuality, and conciseness—improved significantly from 4.42 to 4.58. This successfully crossed the >4.5 target threshold.

The LLM's performance score measured on a scale of 1 to 5 for helpfulness, factuality, and conciseness—improved significantly from 4.42 to 4.58. This successfully crossed the >4.5 target threshold.

Null Search from 9% to 4.1%

The Co-Pilot-led journeys achieved a null search rate of just 4.1%. This is significantly lower (by approximately 5 percentage points) compared to the standard Tata Neu search, which hovered at an 8% to 9% null search rate.

The Co-Pilot-led journeys achieved a null search rate of just 4.1%. This is significantly lower (by approximately 5 percentage points) compared to the standard Tata Neu search, which hovered at an 8% to 9% null search rate.

41.4% AI search adoption

Higher SERP CTR from AI search led journey as compared to 10% regular search SERP CTR

Higher SERP CTR from AI search led journey as compared to 10% regular search SERP CTR

.say hello

© 2026 nirmal tandel. All rights reserved.

nirmal

© 2026 nirmal tandel. All rights reserved.