By Eugenia Nemkova, Marketing Director at Trembit · Last updated: 30 September 2026
To evaluate a WebRTC development partner, test three things: protocol-level real-time expertise (not “we’ve integrated Twilio”), production evidence you can verify, and contract terms that protect your code and your continuity. A generalist web agency and a WebRTC specialist can look identical on a portfolio page. They fail very differently once a product is live, and the difference only shows up in how they answer a handful of specific questions. Below are the 12 questions to ask a WebRTC development company, each with what a strong answer sounds like and what a red-flag answer sounds like, so you can run the evaluation without already being a WebRTC expert yourself.
Key takeaways
- The strongest single signal is whether a vendor can walk you through a specific production problem at the protocol layer (ICE, TURN, the media server), not just list technologies.
- Settle who owns the source code and the infrastructure accounts in writing, before you sign.
- A vendor that can’t show load-test results near your scale, including what broke first, hasn’t run the test.
- “What happens if your engineers leave mid-project?” is the most-skipped question and one of the most consequential.
- Specialist WebRTC firms, generalist agencies and CPaaS platforms each fit different jobs. The right choice depends on what you’re building, not on which sounds most impressive.
Who this is for: CTOs, founders and product leads choosing a vendor for a video or voice product. If you’re not an engineer, run Questions 5–10 yourself and hand Questions 1–4, 11 and 12 to your VP Engineering or a trusted engineer. Each question ends with a one-line “for your product” note in plain business terms. The vendor-type comparison is near the end.
1. Which media server have you built or operated in production, and why that one?
The media server (usually an SFU, a Selective Forwarding Unit) sets your scaling ceiling, a large part of your hosting bill, and what you can do with the media later. A vendor that has actually run one explains the choice through constraints: participants per room, audio versus video, where media is allowed to flow, and what they had to change inside the server.

Strong answer: names a specific stack (mediasoup, Janus, LiveKit, Jitsi, Pion or a custom SFU), the constraint that decided it, what they modified, and what they’d choose differently for your product.
Red flag: a technology list with no trade-off (“we work with all of them”), or an answer that stops at an SDK when you asked about a server.
For your product: this choice decides how many people fit in a call, what hosting costs each month, and whether recording or AI can be added later without a rebuild.
Two examples from our own work. For HearSports, a live sports-commentary app, we chose Janus and tuned it specifically for audio-only forwarding: buffer sizes, jitter compensation and packet-loss concealment set for speech, because listeners notice a 50 ms audio glitch they would never see in a video frame. For Cloudbreak’s in-call medical translation, we extended a Mediasoup SFU with C/C++ modules that intercept audio at the RTP level. That turned a forwarding server into a media-processing engine, and the scaling assumptions that come with pure forwarding had to be re-engineered. Either story tells you more than a logo wall.
2. Can you show me load-test results at a scale close to mine?
WebRTC failures are load-shaped. A room that works with two people can freeze at ten, and a demo proves nothing about the fiftieth participant or the thousandth viewer.
Strong answer: numbers plus method plus the break point. How many concurrent sessions and participants per room, how the load was generated (headless browsers or synthetic clients), which resource hit the ceiling first (SFU CPU, bandwidth, TURN capacity), and what they changed as a result.
Red flag: “it scales horizontally,” “we use Kubernetes,” or a capacity figure with no breaking point. A vendor that never found the limit didn’t look for it.
For your product: without a real load test, your first big customer or event becomes the load test, in front of paying users.
On a live horse-auction platform we built, the load test simulated real auction conditions: hundreds of concurrent bidders, multiple simultaneous video streams and rapid-fire bid sequences. It checked that latency, bid ordering and video-to-bid synchronization held under peak load, before go-live. Notice what was measured: the things that would break the business (a stale bid on a remote bidder’s screen), not only server CPU. Our write-up on bot-based load testing for WebRTC video conferencing covers the method.
3. How do you handle users behind restrictive corporate, hospital or hotel networks?
Some of your users will always sit behind firewalls that block direct peer connections or UDP traffic altogether. WebRTC’s answer is a TURN relay, standardized in IETF RFC 8656: an intermediate server that relays media when a direct path isn’t possible. Whether TURN is set up well decides whether those users can connect at all. It’s also a cost and data-location question, because relayed media runs through servers you pay for.
Strong answer: TURN servers placed in regions near your users; fallback to TCP and TLS on port 443 for networks that block UDP; short-lived credentials; relay usage tracked in production. And a clear testing method, such as forcing relay-only connections (the iceTransportPolicy: "relay" setting in the W3C WebRTC API) to prove the relay path works before users find out it doesn’t.
Red flag: “we use Google’s public STUN server” as the complete answer. STUN helps a client discover its public address; it doesn’t relay anything. Also a red flag: testing only on office Wi-Fi.
For your product: weak TURN setup means some users, often in hospitals and large enterprises, can’t join at all, and they blame your product, not their firewall.
Sometimes compliance dictates the network design. For webPRAX Face2Face, a psychotherapy video platform in Germany, STUN and TURN run on the platform’s own servers, with Germany-only hosting and no third-party media relay. Ask the vendor whether they have ever had to self-host TURN, and why.
4. What happens to call quality on a bad network?
Real users are on trains, patchy LTE and crowded café Wi-Fi. “Real-time video that works” should mean it degrades gracefully and recovers, not that it works on a good connection.
Strong answer: specific mechanisms. Simulcast or SVC (scalable video coding), so the server can send a lower-quality layer to one weak receiver without degrading everyone else. Bandwidth estimation. Protecting audio when bandwidth collapses, because people tolerate a blurry picture far longer than broken speech. Reconnection (ICE restart) when a phone switches networks. And testing under emulated packet loss, jitter and bandwidth caps, not only on a fast office line.
Red flag: “WebRTC handles that automatically.” The browser adapts its own sender; it doesn’t make your server’s forwarding decisions, design your reconnection behaviour or handle a mobile network handover for you.
For your product: most users judge your product by how it behaves on their worst connection, and that is where churn and support tickets come from.
For HearSports we built reconnection and buffer-management logic so that moving from Wi-Fi to cellular mid-match didn’t break sync or interrupt playback. For the horse-auction platform, adaptive bitrate kept the stream stable for bidders watching from barns and fields on cellular connections.
5. Who owns the source code, cloud accounts and domains, in writing?
This is the contract question with the biggest consequences, and the cheapest to settle before signing. If the repositories, cloud accounts, TURN servers, app-store accounts or DNS sit in the vendor’s name, changing vendors later turns into a negotiation.

Strong answer: full IP transfer, written into the contract. Repositories and cloud accounts in your organisation from day one, with the vendor given access rather than ownership. Documentation and knowledge transfer listed as deliverables. Willingness to work under your IP agreement, and a list of the open-source components in the media stack with their licences.
Red flag: vague wording, “our proprietary framework” underneath your product with licence terms that aren’t in the contract, or “we’ll transfer everything at the end.”
For your product: if you don’t own the code and accounts, switching vendors costs time and money at exactly the moment you can least afford it.
A useful benchmark is wording a vendor will commit to publicly. Ours, from our engagement terms, reads: “All source code, documentation, architecture diagrams and intellectual property transfer to you, and we can work under your preferred IP agreement.” Any serious vendor should be able to match that sentence in writing.
6. What happens if the engineers on my project leave mid-build?
WebRTC knowledge tends to concentrate in one or two people. When they leave, the signaling flow, TURN configuration and media-server settings become a black box, and every fix turns into guesswork. A lot of rescue projects start exactly here.
Strong answer: a concrete continuity practice. At least two engineers who know each critical component. Architecture documentation (signaling flow, TURN setup, media-server configuration) as an ongoing deliverable, not a promise. A named replacement process and a realistic onboarding time for a new engineer.
Red flag: “our people don’t leave,” or “we’ll find someone.”
For your product: losing the one engineer who understands your media layer can freeze your roadmap for months.
What good continuity looks like in practice: when Learnster’s corporate learning platform stalled under a previous engineering team, we transferred a full team in one week, so delivery didn’t stop during the handover, and we kept a part-time backup squad so no single person became a point of failure. Seven years on, we still run that engineering programme. Learnster’s engagement is about continuity rather than WebRTC internals; the lesson is continuity discipline, which works the same way on any stack.
When continuity fails on a live video product, the result is usually a rescue. Our WebRTC rescue playbook shows how inherited code gets diagnosed layer by layer, and what a rescue engagement actually diagnoses is laid out on its own page.
Already mid-evaluation and want to hear how a specialist answers these? The last section explains how to put them to our engineers.
7. What compliance work have you actually shipped, not just claimed?
HIPAA, GDPR and Germany’s KBV requirements aren’t badges. They shape the architecture: where media flows, which subprocessors touch it, what gets logged, where it’s hosted. A vendor that adds compliance at the end is quoting you a rebuild.
Strong answer: a specific outcome (a certification, a passed audit, an official listing) and the architecture decisions it forced. Plus readiness to sign the paperwork: a data processing agreement where they handle personal data (GDPR Article 28 requires processing by a processor to be governed by a contract), and a HIPAA Business Associate Agreement where their engineers may access protected health information.
Red flag: “our platform is HIPAA-compliant,” with no answer to who hosts TURN, where recordings live and what gets logged. Or reluctance to sign a DPA or BAA.
For your product: compliance added late means a delayed launch or a partial rebuild before you can sell to regulated customers.
Two examples of what “shipped” means. webPRAX Face2Face, the psychotherapy video platform we built for Healthy Projects GmbH, is certified against the KBV’s video-consultation requirements and appears on the KBV’s register of certified video service providers (list as of 24 August 2026; certificate valid until 13 February 2028). A detail worth testing a vendor on: the KBV doesn’t certify platforms itself. Independent certification bodies check them against Annex 31b of the German physicians’ federal collective agreement (BMV-Ä), and the KBV publishes the list. A vendor that knows this has probably done the work. And Cloudbreak’s in-call translation pipeline was built for HIPAA and GDPR with zero persistence: fixed-size buffers that overwrite themselves, no swap to disk, and an audit log that records that translation happened, not what was said.
8. What does support look like after launch?
A WebRTC product doesn’t stay stable on its own after launch. Browsers change WebRTC behaviour on their own release schedules, and a platform nobody watches degrades quietly until users notice.
Strong answer: named engineers who know the real-time layer, not a generic ticket queue. Severity levels defined around your product (a drop in connection success is critical; a UI glitch isn’t). A clear statement of who gets paged outside business hours and what that cover includes. Who pays for bugs found in production: the vendor fixing its own defects, or you paying for them as change requests. And browser-compatibility testing ahead of new browser releases.
Red flag: “24/7 support” with no names attached, support run by a team that didn’t build the product, or no distinction between defects and new work.
For your product: without real support, a browser update can quietly break calls for some of your users before anyone on your side notices.
For comparison, our published Support & Maintenance model covers production call-quality monitoring and drop-rate tracking, proactive browser-compatibility updates, performance work, and documentation so your team isn’t dependent on us for basic operations. Ask every vendor for their equivalent in writing.
9. Can I talk to a reference client, and ask them about an incident rather than the launch?
References are selected to tell the happy-path story. The useful information is how the vendor behaved when something broke.
Script for the reference call:
- What broke in production, and did you hear about it from the vendor or from your users?
- How long did it take from the first report to a root cause?
- Did the estimate hold? If not, when were you told?
- Is the team you work with now the team you started with?
- Would you hire them again for your next real-time product?
Strong answer: the vendor offers one reference that still works with them and one whose project has ended, and is comfortable with you asking about incidents.
Red flag: references only for the launch story, or a reference who can’t name a single engineer on the project. Public review profiles help here too: read them for patterns across several clients rather than for the star rating.
For your product: how a vendor handled someone else’s outage is the best preview you’ll get of your own.
10. Would you tell me if I don’t need a custom build?
A specialist firm makes money on custom work, so this question tests honesty. CPaaS platforms (communications-platform-as-a-service, such as Twilio, Agora, Vonage or LiveKit Cloud) are often the right answer for an early product: you get working video quickly and run no media infrastructure. A custom build tends to pay off at higher usage, when compliance limits which third parties can touch the media, or when you need to change what happens to the media itself.
Strong answer: before recommending anything, the vendor asks about your expected minutes, participants per session, compliance constraints and roadmap. They’re willing to say “start on a platform, and design so you can move later,” and they can explain what drives the break-even point: usage-based platform pricing on one side, hosting plus engineering on the other.
Red flag: a custom build recommended on the first call, before anyone has asked about your volume.
For your product: the wrong build-or-buy call either spends months of runway on custom work you didn’t need, or locks you into per-minute costs that grow with every user.
A specialist should use managed components where they fit. For Bullpit, a live-streaming platform for investor communities, we used Millicast, a managed WebRTC broadcast service, as the sub-second streaming layer, and put the custom engineering into the product: dual backends, real-time chat and community features.
11. How will you monitor call quality and respond to incidents once we’re live?
“It randomly fails” usually means nobody is collecting data. WebRTC exposes detailed per-connection statistics through getStats(), defined in the W3C WebRTC Statistics specification: packet loss, jitter, round-trip time, dropped video frames and more. If the vendor doesn’t collect them, you’ll be debugging from user complaints.
Strong answer: client-side stats collected per session and sent to a backend; server-side metrics (media-server load, TURN usage); dashboards and alerts on connection success and call drop rate. And a walk-through of the last real incident: how it was detected, the root cause, the fix and what changed afterwards.
Red flag: “we monitor server uptime.” A server can be up while calls fail. Also a red flag: no incident they can describe.
For your product: without call-quality data, you learn about problems from customer complaints and churn instead of from a dashboard.
On a government video-conferencing platform we built on Janus, monitoring was part of the design from the start: Prometheus metrics for media-server health and connection quality, Grafana dashboards with alerting, and Elasticsearch for searchable session and security logs. webPRAX Face2Face runs Grafana and Kibana for uptime and incident response. For the most common production failures that monitoring catches, see why WebRTC projects fail in production.
12. How would you add AI features (transcription, translation, voice agents) to this stack later?
Few vendors get asked this yet, but it’s worth asking even if AI isn’t in scope. The architecture decision that matters (where the audio is tapped) gets made now. Audio can be captured on the client, inside the media server, or by a bot that joins the room as a participant. Each option has different consequences for latency, cost and compliance.
Strong answer: where they would tap the audio and why; a latency budget per stage (speech recognition, the model, speech synthesis); what happens to the audio from a compliance point of view (does it leave your infrastructure, and is the AI provider covered by a BAA or DPA?); and, for voice agents, how they handle turn-taking and interruptions.
Red flag: “we’ll call the OpenAI API,” with no answer about where the audio goes or how long each step takes.
For your product: where the audio is tapped today decides whether adding AI next year is a feature release or a re-architecture.
The answer should draw on something shipped. For Cloudbreak we built speech recognition, machine translation and speech synthesis into the live call itself, intercepting audio inside a custom Mediasoup SFU at the RTP level. The result was dubbed audio and live subtitles at sub-second latency, plus a confidence layer that shows the original text instead of dubbing a medical phrase the model isn’t sure about. It was one combined delivery, not two systems bolted together.
Specialist WebRTC firm, generalist agency, or CPaaS platform: which do you need?
The 12 questions work on any vendor, but you’ll weigh the answers differently depending on which type you need.

| Option | Strongest when | Trade-off to accept | Examples |
|---|---|---|---|
| Specialist WebRTC firm | Real-time video or voice is the product: custom architecture, scale, compliance constraints, changes to the media itself, or rescuing a failing build | Fewer firms to choose from; their incentive leans toward custom work (use Question 10) | Trembit, Fora Soft, WebRTC.ventures |
| Generalist software agency | Video is a small feature inside a broader app, usually built on a CPaaS SDK | Protocol-level problems (TURN, SFU scaling, codec negotiation) can sit beyond their depth once usage grows | Most full-service agencies |
| CPaaS platform | You want working video fast and can accept the platform’s architecture | Usage-based pricing that grows with your minutes; less control over the media path; the vendor becomes a subprocessor for compliance | Twilio, Agora, Vonage, LiveKit Cloud |
These combine. A common path is a CPaaS platform plus a generalist agency for the first version, then a specialist once usage, compliance or quality problems outgrow the platform. Once you’ve narrowed the type and have a shortlist, you can compare specific WebRTC development companies side by side.
Want to put these questions to us?
We’d rather be tested than trusted. Bring all 12 questions to a free 30-minute technical call with a Trembit engineer. No deck, no pitch. We’ll answer each one with the same specifics this guide asks of everyone, including Question 10, whether you need a custom build at all. If you do, that’s the work our WebRTC development team does every day. If you don’t, you’ll leave with a clearer brief for whichever vendor you pick.
FAQ
How many WebRTC vendors should I shortlist? Two to four. That’s enough to compare answers to the same questions side by side, and few enough that the evaluation doesn’t stall. If you start with a longer list, use Questions 1, 5 and 7 as a first screen.
Should compliance questions come before technical ones? If compliance is a hard requirement (HIPAA, GDPR with EU-only hosting, KBV), ask first. It can eliminate candidates before you spend time on technical vetting, and it changes how you judge every later answer, from TURN hosting to AI features.
How long should a WebRTC vendor evaluation take? For most teams, a few weeks: one technical call per vendor built around these questions, a reference call per finalist, and a contract review focused on IP, accounts and continuity. Allow longer if a compliance review is involved.
Is fixed-price or time-and-materials better for WebRTC work? It depends on how certain the scope is. Fixed price suits a clearly specified build, such as an MVP or a defined feature, where both sides can agree on the deliverables up front. Time and materials suits work that starts with diagnosis or will change as you learn, such as a rescue or scaling an existing platform, where the exact fix only becomes clear once engineers are inside the code. Long-lived products with an evolving roadmap often move to a dedicated team.
What if a vendor can’t answer several of these questions? That’s the signal. One weak answer can be an honest gap; a vendor that is vague on media servers, load testing and TURN strategy together hasn’t operated real-time video in production. A vendor that says “we haven’t done that, here’s how we’d find out” is more trustworthy than one that improvises.