8 min read

Lessons from Our First Five Early-Access Partners

Lessons from intella's first five early-access contact center partners

Customer testimonials are one measure; production failures are another. We opened intella's early-access program in late 2024 with five contact center partners across GCC banking and telecom. Eight months on, those deployments had exposed more about our mistaken assumptions around Arabic call data than testimonials or case study figures could.

This is an account of those surprises: what we built, what we rebuilt, and what the data revealed unexpectedly. We are sharing it to give teams assessing Arabic voice AI tools useful context, and to show why early-stage development is better understood through its working details than its polished presentation.

The five partners

We chose partners with deliberately different profiles so the product would face varied conditions. Two were retail banking contact centers: one in Riyadh handling several Gulf Arabic dialects, and one in Kuwait serving a mainly Khaleeji-speaking customer base. Two were telecom operators: one Saudi mobile operator with heavy prepaid volume, and one serving a Levant market where Levantine was the dominant register. The fifth was an insurance carrier handling inbound claims, giving us longer calls, a more formal register, and fewer call-and-transfer patterns than banking and telecom.

The industry and dialect mix was intentional. We needed to separate capabilities that generalized from those that worked only because the input had been narrowly tuned.

The dialect distribution surprise

Our first model treated dialect as a classification task: identify the dialect at the beginning of a call, then apply the matching model weights. That approach performed reasonably on curated audio. In production, it failed because our assumption about how dialect is distributed throughout a call was incorrect.

Production conversations do not keep one stable dialect. An agent in Riyadh speaking with a customer who moved from Egypt may shift toward that customer's register. A Levantine caller reaching a Gulf operator may adopt Gulf vocabulary for domain terms because the agent uses it. One call can contain stretches in two dialect registers, so a classification run once at the start can label the latter half incorrectly.

We replaced one-time classification with continuous dialect re-evaluation. It increased processing time, yet raised transcript quality by a meaningful margin on the most ambiguous calls. Production evidence made the change possible; our curated development samples had not included the same register-switching behavior.

Code-switching exceeded our estimate

We expected code-switching and had trained on datasets containing some mixed Arabic-English speech. What we misjudged was the concentration of English financial and technical vocabulary in Gulf Arabic banking calls.

At the Riyadh banking contact center, about 15 to 20 percent of tokens in an average call were English terms inside Arabic sentences. For ordinary transaction terms such as "transfer," "statement," and "balance," English appeared as often as Arabic, and sometimes more often. English was nearly universal for digital banking product names and feature labels.

That density meant a system focused mainly on Arabic tokens would consistently miss or mishandle a substantial share of each call's useful content. We added a code-switching layer to process the boundaries between languages instead of treating English tokens as noise. On calls with heavy English usage, it delivered our largest single quality improvement during early access.

The reporting format nearly shipped wrong

At the program's start, we sent pilot partners weekly reports designed around our internal view of the data: call clusters ordered by churn risk score, sample transcript excerpts, and a confidence indicator. Feedback from two partners during the first month matched: the reports were technically interesting, but not actionable for the people receiving them.

We had designed for a data analyst who would turn clusters into business decisions. The actual recipients were customer experience directors and contact center operations managers. They needed direction, not only a description of the data. To an analyst, "cancellation intent, topic: fee dispute, high confidence" may be useful. To an operations manager, "12 customers this week expressed cancellation intent specifically referencing the monthly account maintenance fee, most called more than once, here is what the call flow looked like" supports action.

Following those two partners' feedback, we rebuilt the templates and placed a plain-language summary above the detailed cluster data. The revised format drew much stronger participation in weekly pilot reviews. A report that people read and act on is the product, rather than the analytics underneath it.

Batch processing was right first, but pressure arrived sooner

We started with batch processing: calls ran nightly or weekly, producing a periodic report. We intentionally left real-time processing out of the first version because it would have delayed launch considerably, and we did not know whether the targeted use cases required it.

The telecom pilots showed that demand for lower latency would arrive earlier in the product lifecycle than expected. By week four, one telco partner was asking about same-day processing for a use case we had not fully considered: responding to service outages. When a regional outage brings hundreds of customer calls in during the same hour, the operator wants to understand those conversations within hours, not the next day. It is not exactly churn detection, but it is a genuine call-intelligence use case that batch processing could not support.

We are not suggesting batch processing was wrong for launch. It suited the launch. We moved real-time processing earlier in the development timeline because of what the pilots told us, so the early-access architecture reflects observed needs rather than only our starting product assumptions.

What we missed in churn vocabulary

Our first churn layer focused on explicit cancellation and dissatisfaction language: statements about closing an account, leaving a carrier, or cancelling a service. Those expressions are predictive when present, but they occur in only a fraction of calls from customers who later churn.

Banking pilot data pointed to an earlier and often stronger signal: repeated frustration about the same issue across several contacts before anyone uses explicit cancellation language. A customer who calls three times about one unresolved problem is a stronger churn indicator than someone who says "I want to cancel" once. The repeated pattern shows the bank had multiple chances to resolve the issue and did not.

That finding changed the detection logic considerably. The current intella version treats repeated contacts around a specific topic as a primary signal, while explicit cancellation language confirms the signal instead of triggering it. Data drove the change, not our initial view of how churn appears in Arabic recordings.

What remains unresolved

We are not suggesting we have solved Arabic call intelligence. Five partners and eight months of data provide useful evidence, but not complete coverage. Some dialect variants remain underrepresented. Some industry vocabularies are not fully mapped. Certain interactions, including collection calls and complex loan applications, still need better topic classification.

The early-access program supplied what a young product needs most: production data from contact centers that challenged our starting assumptions. The partners who contributed their time shaped the product for those who come next, which is the purpose of the program. Before the next deployment, test dialect shifts, code-switching, repeat-contact signals, and report actions against production calls before fixing the architecture.