Key Takeaways
- Ringg handles more than 7 million connected calls a month and resolves up to 65% of routine customer requests without human agents, serving Policybazaar, Practo, and Groww.
- A routing layer picks between GPT-4.1, GPT-5.6 Luna, Terra, and Sol per request, while production evals and live-traffic canaries tune quality, latency, and unit economics at roughly 90% lower cost than GPT-4.1.
- Documented outcomes include an 88% faster response time at Policybazaar, 85% first-call resolution at Practo, and 72% self-service resolution at Groww, shifting the core metric from call volume to completed outcomes.
Table of Contents
Ringg’s 7 Million Call Automation Benchmark
A voice and chat agent platform built for large consumer operations in India now handles a monthly volume beyond 7 million connected calls while resolving up to 65 percent of routine customer requests without human intervention. That is the headline finding from a new technical profile published by OpenAI, which details how Ringg uses the GPT-5.6 model family across voice, chat, WhatsApp, and web at roughly 90 percent lower cost than GPT-4.1 on selected workloads.
Ringg’s average customer satisfaction score sits at 4.8, and its deployments include Policybazaar, Practo, and Groww. The underlying shift is not just speed; it is about moving customer operations from headcount scaling to outcome-based automation.
Model quality is only part of the equation. We also need low latency, reliable tool use, strong instruction following, and economics that work at scale. OpenAI gave us the balance we needed.
— Siddharth Tripathi, Co-founder, Ringg
Routing, Evals, and the Model Stack Behind the Resolution Rate
In its model selection process, Ringg compared OpenAI with alternatives across conversational quality, latency, instruction following, tool use, multilingual performance, reliability, and cost. The company concluded that the OpenAI stack delivered the most consistent balance for production workloads.
Ringg’s architecture does not depend on a single model. Each request passes through a routing layer that selects a model based on latency, instruction-following needs, tool use, and price-performance.
The platform coordinates multi-step workflows across CRMs, ticketing systems, payment tools, scheduling APIs, and internal systems. When an interaction requires escalation, the system hands off the conversation to a human with a summary intact.
- GPT-4.1: carries the majority of live voice and chat traffic.
- GPT-5.6 Luna: remains in production for real-time interactions that require stronger performance, lower latency, or better unit economics.
- GPT-5.6 Terra: powers post-call analysis such as summaries and sentiment classification.
- GPT-5.6 Sol: supports evaluation, prompt improvement, and model-as-judge workflows.
This routing approach matters because a customer service call is not a single inference task. The orchestration layer fuses live customer input with the agent’s instructions, conversation records, account data, relevant knowledge-base material, and tool access before a model is chosen.
For longer conversations, the system compresses the dialogue into a structured summary once the context nears roughly 80,000 tokens. The exchange can then continue with key information preserved instead of resending the entire transcript each turn.
Production evals, not just model benchmarks
Before a model reaches production, Ringg evaluates it against recorded conversations and simulated customer journeys. The evaluation platform identifies failure points and suggests prompt changes, creating a loop that sharpens both quality and unit economics.
In one test, GPT-5.6 Terra outperformed Gemini 2.5 Flash on post-call analysis, including summary and sentiment accuracy. Terra also held accuracy on regional languages, with up to 97 percent accuracy on common languages in the markets Ringg serves.
After passing offline tests, a model is exposed to a small slice of live traffic before the rollout broadens. Once live, the routing system watches regional latency and endpoint health, rerouting traffic if a service becomes unreachable or exceeds a response-time threshold.
Customer results across insurance, healthcare, and investing
Policybazaar, a major online insurance platform in India, connects more than 57,000 customer requests through Ringg, with 67 percent of calls handled without human intervention. Its average response time dropped from between 8 and 12 minutes to under 60 seconds, an improvement of roughly 88 percent.
Practo, a global healthcare platform, used Ringg to support appointment booking and service completion. It reported an 85 percent first-call resolution rate, sub-three-second response times, and a 70 percent reduction in operating costs compared with human-led workflows.
Groww resolves 72 percent of inbound IPO, futures, and options queries entirely through self-service, with an average handling time of two minutes.
Ringg is also building browser agents for onboarding, Know Your Customer checks, IT troubleshooting, incident support, and claims processing. A planned context layer would let a customer start a request on voice, continue on WhatsApp, and finish in a browser without repeating details.
Why the Customer Service Market Is Splitting Between Automation and Trust
According to Zendesk research, 86 percent of CX leaders using AI and automation report significant cost savings, while Zendesk’s own AI agents can automate up to 80 percent of customer interactions.
Gartner projects that conversational AI will reduce contact center labor costs by $80 billion in 2026. By 2029, agentic AI could autonomously resolve 80 percent of common customer service issues without human intervention.
Yet adoption is not the same as acceptance. SurveyMonkey research from late 2025 found that 79 percent of U.S. consumers still prefer speaking with a human agent, and 84 percent believe human agents provide more accurate support.
The tension is not about whether AI can deflect calls. It is about whether the business case for automation can hold once consumers demand clearer escalation paths, explainability, and continuity across channels.
Ringg’s customer examples suggest the efficiency story extends beyond support tickets into revenue-adjacent workflows: appointment bookings, policy connections, and investment self-service. That autonomy raises the bar for fallback behavior and human handoff quality.
The hard part is that human escalation remains necessary for sensitive, complex, or high-risk cases. The real differentiator may not be the highest automation rate, but the ability to preserve trust in the interactions that still need a person.
The Next Metric Is the Completed Outcome, Not the Call Count
Ringg’s shift from call volume to completed business outcomes — insurance policies connected, appointments booked, investment queries resolved — marks the most important change in customer service economics. For AI teams, the new benchmark is no longer whether an agent can deflect a call, but whether it can finish a fragmented, multi-step task without losing the customer’s context. For teams building AI agent workflows that need to scale without breaking consumer trust, programmatic SEO and AI automation is how Andres SEO Expert approaches the challenge — get in touch.
Frequently Asked Questions
What is Ringg’s 7 million call automation benchmark?
Ringg is a voice and chat agent platform for large consumer operations in India that handles over 7 million connected calls per month while resolving up to 65 percent of routine customer requests without human intervention. It uses the GPT-5.6 model family across voice, chat, WhatsApp, and web at roughly 90 percent lower cost than GPT-4.1 on selected workloads.
How does Ringg use GPT-5.6 models in production?
Ringg routes requests across multiple models: GPT-4.1 carries most live voice and chat traffic, GPT-5.6 Luna handles real-time interactions needing stronger performance or lower latency, GPT-5.6 Terra powers post-call analysis like summaries and sentiment, and GPT-5.6 Sol supports evaluation and model-as-judge workflows.
How does Ringg’s routing and evaluation process work?
Each request passes through a routing layer that selects a model based on latency, instruction following, tool use, and price-performance. Before production, models are evaluated on recorded conversations and simulated journeys. After passing offline tests, they get a small slice of live traffic, then broader rollout, with live monitoring and rerouting if latency or endpoint health degrades.
What customer results did Policybazaar, Practo, and Groww achieve with Ringg?
Policybazaar connects more than 57,000 customer requests through Ringg with 67 percent of calls handled without human intervention and response time dropping from 8-12 minutes to under 60 seconds. Practo reported 85 percent first-call resolution, sub-three-second response times, and 70 percent lower operating costs. Groww resolves 72 percent of inbound IPO, futures, and options queries through self-service with a two-minute average handling time.
Why is the customer service market splitting between automation and trust?
Zendesk research says 86 percent of CX leaders using AI and automation report significant cost savings, and Gartner projects conversational AI will cut contact center labor costs by $80 billion in 2026. However, SurveyMonkey found 79 percent of U.S. consumers still prefer a human agent, and 84 percent believe humans provide more accurate support. The challenge is preserving trust and clear escalation paths.
What is the next metric for customer service AI after call volume?
The next metric is the completed business outcome, such as insurance policies connected, appointments booked, or investment queries resolved. The key question is whether an AI agent can finish a fragmented, multi-step task without losing the customer’s context, not just whether it can deflect a call.
How does Ringg handle long conversations and human escalation?
When context nears roughly 80,000 tokens, Ringg compresses the dialogue into a structured summary so the exchange can continue without resending the full transcript. If an interaction requires escalation, the system hands off to a human with a summary intact.
