Do I need special permissions for voice features in my app?
The question usually arrives late. A product is weeks from launch, voice search or a voice note feature has been added to the roadmap, and someone in the room asks whether the app actually needs permission to use the microphone. The answer is yes, always, on every platform, with no exceptions. But the more useful question is what kind of permission you need, when to ask for it, and how the moment you ask shapes whether users say yes or quietly close the app and never return.
The question is not whether you need permission, but when and how you ask for it.
We have seen this decision made well and badly. On a peer-to-peer currency exchange product we built, the core transfer mechanism worked as intended, but we had not factored in anti-money laundering requirements. Apple flagged it, and we had to retrofit KYC checks, enhanced transfer security, and hard limits on transfers between parties. The rework cost time and budget that would not have been needed if compliance had been part of the original design. Voice permissions carry similar risk. Build first, ask questions later, and you pay twice.
This article covers what the permissions actually are, what Apple and Google require, and why the gap between technical compliance and genuine user trust is where most products quietly lose their users.
What permissions does a voice feature actually require?
Any feature that captures audio from a device microphone requires explicit permission from the user. This applies whether the feature is voice search, voice messaging, audio recording, speech-to-text input, or a voice-activated assistant. The permission is not implied by the user downloading the app or agreeing to general terms. It must be requested at the point of use, and the user must actively grant it.
Microphone access
Microphone access is the baseline requirement. On both iOS and Android, this is a runtime permission, meaning the system dialogue appears during the session when the feature is first triggered. The app cannot access the microphone in the background without additional justification, and that justification must be declared in advance.
Speech recognition and processing
If the app converts speech to text, a second consideration arises. On iOS, apps using the Speech Recognition framework must declare this separately and explain to users that audio may be sent to Apple's servers for processing. Where a third-party speech API is involved, such as Google Cloud Speech or a similar service, the audio is being transmitted off-device, which carries its own data handling obligations under GDPR and equivalent frameworks. Users are rarely told this clearly, and the absence of that clarity is a trust problem as much as a compliance one.
Platform rules: what Apple and Google mandate
Apple and Google both require microphone access to be declared before the app is submitted for review. On iOS, the NSMicrophoneUsageDescription key must appear in the app's Info.plist file, and the string value must describe why the microphone is needed in plain language. Apple's reviewers read this string, and vague entries such as "for app functionality" are grounds for rejection. The description must match the actual use.
On Android, the RECORD_AUDIO permission must be declared in the app's manifest. From Android 6.0 onwards, this is a runtime permission, so the system dialogue appears in-app rather than at install time. Google Play also reviews permission declarations and will flag apps that request permissions inconsistent with their stated purpose.
We saw the cost of missing this kind of review on the currency exchange product. Apple's review process identified a gap the development team had not anticipated, and the retrofit took considerably longer than the original build. The same pattern applies to voice features. A missing or vague NSMicrophoneUsageDescription stops the submission. A feature that processes speech through a third-party API without disclosing it can trigger a rejection at review or, worse, a removal after launch.
Both platforms permit users to revoke permissions at any time through device settings, so the app must handle a declined or revoked microphone permission gracefully, without crashing or leaving the user stuck on a non-functional screen.
Start your app project the right way
We deliver the complete blueprint before a line of code is written. User research, psychology-driven design and full technical specifications. You choose who builds it.
The compliance floor is not the same as user trust
Meeting Apple's and Google's requirements gets the app into the store. It does not persuade anyone to tap Allow when the system dialogue appears. Those are two different problems, and confusing them is where a lot of voice features fall apart in practice.
The system permission dialogue is a fixed format. The app cannot alter its wording, layout, or the two options it presents. What the app can control is everything that happens in the moments before that dialogue appears: what the user has already experienced, how the request has been framed, and whether they understand what they are agreeing to.
Simon's framework for thinking about trust puts it plainly. Trust becomes relevant when a product asks something of the user. Browsing a content feed carries no real trust stakes. Granting microphone access to an app the user has been using for thirty seconds carries significant ones. The user is being asked to allow a piece of software to listen through their device, and the product has usually done very little to deserve that level of confidence.
Trust stakes rise the moment a product asks something of the user, and microphone access asks a great deal.
Framing the request well, explaining the benefit clearly, and sequencing the ask at a moment when the user already sees value in the feature are all things the product controls. They are framing and tone-of-voice decisions, and they matter more than the permission string in the Info.plist.
Why asking too early destroys conversion
The moment a new user opens an app, they are still forming their first impression. They do not yet know whether the app is worth their time, whether the product does what it promises, or whether the people behind it are trustworthy. Asking for microphone access in that state is asking for a high-trust commitment before any trust has been built.
According to Localytics, apps that request permissions in the first session see up to 60% lower opt-in rates compared to apps that wait until the user understands the value being unlocked. That gap is primarily a sequencing problem.
We worked on a map-based fitness social network where users could connect with others for shared runs and cycle rides. A significant proportion of users were dropping off at the point where they were asked to share their precise location. The drop-off happened because the location request arrived before users had any opportunity to learn about the other person. We were asking for a high-trust action before the product had given users any reason to extend that trust.
Voice features face the same dynamic. A user who has just downloaded a recipe app and tapped on their first recipe has no frame of reference for why a microphone request is appearing. The app has not yet demonstrated that the voice feature makes their experience meaningfully better. The ask arrives without context, and a significant proportion of users decline it, or worse, delete the app entirely.
Delay the microphone permission request until the user has already experienced the feature's value, or can see clearly what they will gain by enabling it. A pre-permission screen explaining the benefit in plain language, placed immediately before the system dialogue, dramatically increases accept rates.
Sequencing the request: earning the ask before making it
The fix to early drop-off is sequencing. The microphone request should arrive at the moment a user has a clear, immediate reason to want the feature. That moment is product-specific, but the principle is consistent: the user should understand what they gain before the system dialogue appears.
On the fitness social network, we restructured the flow so users could view profiles and exchange messages before being asked to share any location data. Once a conversation had formed and both parties had decided they wanted to meet up, the location-sharing step felt like a natural next move rather than an intrusive demand. The feature was the same. The sequencing changed.
Voice features benefit from the same logic. A voice search capability in a recipe app works better when the user has already typed a search, seen results, and understands how the app works. At that point, showing a screen that says "searching by voice means you can keep your hands free while you cook" gives the microphone request a job to do. The user understands the exchange they are making.
The pre-permission screen
A pre-permission screen sits between the user's action and the system dialogue. It is entirely within the app's control, and it is where the framing work happens. A good pre-permission screen names the specific benefit, reassures the user about what the app will and will not do with the audio, and presents the system dialogue as the natural next step rather than a surprise.
Asking, not demanding
The language of the pre-permission screen matters. Phrasing the request as a question rather than a statement shifts the psychological dynamic. "Can we use your microphone so you can search hands-free?" reads differently to "Microphone access required." The information is the same. The user's sense of control is not. Psychological buy-in increases when people feel they are choosing rather than complying, and that feeling translates directly into higher accept rates and better long-term retention.
When retrofitting compliance costs you twice
Building voice features without thinking through the permission flow is one version of this problem. Building them without thinking through the downstream data handling obligations is another, and the second tends to cost more.
On the dating app project focused on verified profiles and preventing bots, the client chose to skip discovery around the messaging component and focus purely on onboarding. That decision meant a generic messaging feature was built that contradicted the core product premise, allowing automated and fake messages that undermined all the verification work done during onboarding. The mismatch between the rigorously verified onboarding and the permissive messaging system meant the entire messaging section had to be rewritten, costing approximately £15,000 in additional budget and around two months of extra work.
Voice features carry comparable risk. If an app captures audio and transmits it to a third-party speech API, that data flow needs to be declared in the privacy policy and, in some jurisdictions, in the consent flow. If the app stores voice recordings, retention periods and deletion rights under GDPR apply. On the anonymous messaging app we worked on, we had to balance GDPR's right to deletion against the legal need to retain data in case of criminal investigation. We resolved it with a data retention policy of around six months, so that if a user deleted their account after sending harmful content, the data was not immediately wiped. Voice data carries the same kind of tension.
Before building a voice feature, map the full data journey: what is captured, where it is processed, how long it is retained, and whether the user can request deletion. Doing this at the design stage costs an afternoon. Doing it after launch, under pressure from a platform review or a regulatory query, costs considerably more.
How to frame the microphone request so users say yes
The system dialogue is fixed, but everything before it is a design decision. The goal is to reach that dialogue at a moment when the user already wants to say yes.
The pre-permission screen should do three things. It should name the specific benefit the user gets from enabling the microphone, not the generic benefit of the feature category. It should address the most obvious concern directly, typically something like "we only use your microphone when you tap the voice button." And it should make the action feel like a small, natural step rather than a significant commitment.
Language that gives users control
Framing the request as a question rather than an instruction changes the emotional register. "Is it okay if we access your microphone for voice search?" is not meaningfully different in technical outcome from a statement, but the user experiences it differently. They feel the product is asking rather than taking, and that feeling increases the likelihood of a yes. Apps that push notification or access requests without this kind of framing consistently see acceptance rates below 15%, according to Delon Apps.
What to avoid
Avoid framing the request around the app's needs. "We need microphone access to provide this service" centres the product, not the user. Avoid vague reassurances like "your privacy is important to us, " which carry no specific information and are widely ignored. And avoid placing the microphone request alongside three or four other permission requests in a single screen. Stacking permissions in one burst signals that the app is taking rather than asking, and the decline rate climbs sharply.
Test the pre-permission screen with real users before launch. A five-person moderated session will reveal whether the language lands as intended, and adjusting copy costs nothing compared to the cost of a poor opt-in rate at scale.
Conclusion
Voice features require microphone permission on every platform, declared in advance and requested at runtime. That is the technical floor. But the gap between meeting that floor and actually getting users to grant access is where the real design work sits, and it is a framing problem rather than a technical one.
The currency exchange product taught us that compliance gaps found after launch cost significantly more to fix than compliance gaps identified before a line of code is written. The dating app's messaging rewrite, at £15,000 and two months of extra work, reinforced the same point. The fitness social network showed that rearranging the sequence of trust-building interactions, without changing the underlying feature, moves the conversion numbers substantially.
Voice permissions are a high-trust ask. The user is being asked to let the app listen through their device, and the product has usually done very little to earn that level of confidence before the dialogue appears. Sequencing the ask so it arrives when the user already understands the benefit, framing it as a question rather than a demand, and handling the data responsibly on the other side of the yes are all decisions that pay for themselves many times over.
If you are building a voice feature and want to think through the permission flow, the framing, and the data handling before the first build sprint begins, let's talk about your voice feature.
Frequently Asked Questions
Yes, without exception, on every platform. Whether your feature is voice search, voice messaging, or speech-to-text input, the user must actively grant microphone access before the app can use it. Downloading the app or accepting general terms does not count as permission.
The system dialogue should appear at the point of use, meaning when the user first triggers the voice feature during a session. Asking too early, before the user understands why the microphone is needed, significantly reduces the likelihood they will agree.
Apple requires the NSMicrophoneUsageDescription key to be declared in the app's Info.plist file before submission. The description must explain clearly why the microphone is needed, as vague entries such as 'for app functionality' are grounds for rejection by Apple's reviewers.
The RECORD_AUDIO permission must be declared in the app's manifest. From Android 6.0 onwards, this triggers a runtime dialogue within the app rather than appearing at install time, and Google Play will flag apps that request permissions inconsistent with their stated purpose.
On iOS, yes. Apps using the Speech Recognition framework must declare this separately and inform users that audio may be sent to Apple's servers for processing. If a third-party API such as Google Cloud Speech is involved, additional data handling obligations apply under GDPR and similar frameworks.
If your app transmits audio to a third-party service for processing, that data transfer carries obligations under GDPR and equivalent regulations. Users should be clearly informed that their audio leaves the device, as failing to communicate this is both a compliance risk and a trust problem.
Retrofitting permissions and compliance requirements late in development costs significantly more time and budget than building them in from the start. Platform reviewers may reject the app, and rushed permission flows tend to damage user trust at a critical moment.
The moment and manner in which you request microphone access directly affects user consent rates. Users who do not understand why the permission is needed are far more likely to decline or abandon the app entirely, so clear context and timing are as important as technical compliance.