kmail.at
← learning

gradio · difficulty ◆◆

gr.Microphone — a record-only audio component in one line

The smallest component in Gradio that isn't pretending to be flexible — its entire body is 78 lines that set one value and forward 26 parameters.

Pass sources=["upload"] to gr.Microphone and nothing happens. No error, no warning, no upload box — the argument is accepted, forwarded, and then overwritten before it ever reaches the parent class.

2026-10-04 · 5 min read

$ gr.Microphone()

What it does

gr.Microphone is gr.Audio with the choice taken away. Its whole implementation lives in gradio/templates.py as `class Microphone(components.Audio)` with a docstring that reads exactly `Sets: sources=["microphone"]` — no __init__ body beyond one sources assignment and a super() call forwarding the other 26 parameters. As an input it hands your function a (sample_rate, int16 ndarray) tuple by default (type="numpy") or a path inside the Gradio cache (type="filepath"). On top of the shared Audio events it carries the recorder ones: start_recording, pause_recording, stop_recording, plus .stream() for chunked live input. Its format default is "wav" — gr.Audio's is None.

Why it matters

An upload zone is a liability in a demo that only makes sense with a voice sample: users drop a screenshot, an .ogg their browser cannot play, or nothing at all, and you have to detect and explain each case. Microphone removes the decision — one record button, one value shape. Because it is the same class underneath, nothing else changes: waveform_options, subtitles, buttons, playback_position, streaming and both preprocess paths all behave identically to gr.Audio. The one thing you should know before reaching for it is that the recorder events fire once per take, not per chunk, which is what makes them the right place to put transcription, scoring or a save step.

Example

$ import gradio as gr
import numpy as np


def clip_report(audio):              # one component -> one argument
    if audio is None:                # user cleared the take
        return {"status": "no clip"}
    sr, data = audio                 # type='numpy' default: (int rate, int16 array)
    mono = data.mean(axis=1) if data.ndim > 1 else data
    return {
        "sample_rate": int(sr),
        "frames": int(data.shape[0]),
        "seconds": round(data.shape[0] / sr, 3),
        "channels": int(data.shape[1]) if data.ndim > 1 else 1,
        "peak": int(np.abs(mono).max()),
        "clipped": bool(np.abs(mono).max() >= 32767),
    }


with gr.Blocks(title="Voicemail QA") as demo:
    gr.Markdown("## Voicemail QA")
    mic = gr.Microphone(label="hold to record")
    out = gr.JSON(label="what your handler gets")
    mic.stop_recording(clip_report, mic, out, api_name="report")
    mic.clear(lambda: None, None, out, api_name="reset")

demo.launch(prevent_thread_lock=True, server_port=7921)
* Running on local URL:  http://127.0.0.1:7921
* To create a public link, set `share=True` in `launch()`.

Renders: a 'Voicemail QA' heading, an audio widget labeled 'hold to record' that has ONLY the microphone tab (no upload dropzone — the served /config shows sources=['microphone'], format='wav', streaming=False), and a JSON panel labeled 'what your handler gets'. The served /config lists two dependencies: [2,'stop_recording'] api_name='report' and [2,'clear'] api_name='reset'. Driving /report with gradio_client and a real 1.5 s 220 Hz WAV (24000 frames @ 16 kHz) returned exactly {'sample_rate': 16000, 'frames': 24000, 'seconds': 1.5, 'channels': 1, 'peak': 19660, 'clipped': False}.

Verified on gradio 6.27.0 — that dict is the real event payload from the live app. The clip was generated with stdlib wave at 60% amplitude, so peak 19660 is the actual sample maximum, not a rounded guess.

$ import gradio as gr


def echo_path(p):
    import os
    if p is None:
        return "-"
    return "%s | %d bytes" % (os.path.basename(str(p)), os.path.getsize(str(p)))


with gr.Blocks(title="wav-or-not") as demo:
    gr.Markdown("## default format: Microphone vs Audio")
    with gr.Row():
        mic = gr.Microphone(type="filepath", label="gr.Microphone (format default 'wav')")
        aud = gr.Audio(type="filepath", label="gr.Audio (format default None)")
    a = gr.Textbox(label="mic got")
    b = gr.Textbox(label="audio got")
    mic.change(echo_path, mic, a, api_name="mic_path")
    aud.change(echo_path, aud, b, api_name="audio_path")

demo.launch(prevent_thread_lock=True, server_port=7925)
* Running on local URL:  http://127.0.0.1:7925
* To create a public link, set `share=True` in `launch()`.

Renders: a Row with two filepath-mode audio components side by side (the Microphone one still exposes only the mic tab) and two Textboxes below them. Calling the live endpoints with the SAME 4977-byte mp3: /mic_path returned 'audio.wav | 48044 bytes' and /audio_path returned 'audio.mp3 | 4977 bytes'. Both handlers received a file, but only one received the file they uploaded.

This is the practical cost of the wav default: format='wav' on the Microphone made preprocess transcode the mp3 to a 48 kB wav before the handler saw it, while gr.Audio passed the original bytes through untouched. Set format=None on a Microphone to opt out — verified, the attribute then reads None.

$ import gradio as gr
import numpy as np


def level(audio):                    # ONE argument — the (sample_rate, chunk) tuple
    if audio is None:
        return "idle"
    sr, chunk = audio
    if chunk is None or len(chunk) == 0:
        return "silence"
    rms = float(np.sqrt((np.asarray(chunk, dtype=np.float64) ** 2).mean()))
    return "%.1f dBFS" % (20 * np.log10(max(rms, 1) / 32768))


with gr.Blocks(title="Live dictation meter") as demo:
    gr.Markdown("## Live dictation meter")
    mic = gr.Microphone(streaming=True, label="live mic")
    lvl = gr.Textbox(label="level")
    log = gr.Textbox(label="recorder log", lines=4)
    mic.stream(level, mic, lvl, api_name="level")
    mic.start_recording(lambda: "recording started", None, log, api_name="armed")
    mic.stop_recording(lambda: "recording stopped", None, log, api_name="stopped")

demo.launch(prevent_thread_lock=True, server_port=7922)
* Running on local URL:  http://127.0.0.1:7922
* To create a public link, set `share=True` in `launch()`.

Renders: a mic-only audio widget labeled 'live mic' with streaming enabled (served /config: sources=['microphone'], streaming=True, format='wav') and two Textboxes. The /config dependencies show the difference in wiring: the stream dependency is ([[2,'stream']], 'level', connection='stream', stream_every=0.5) while the two recorder events are plain sse dependencies named 'armed' and 'stopped'. Calling level() directly with a real 0.5 s 220 Hz chunk (60% amplitude) returned '-7.4 dBFS'; level(None) returned 'idle'; level((16000, np.array([], dtype=np.int16))) returned 'silence'.

The wrong signature is the classic trap: wiring `def bad(s, c)` to a single Microphone logs 'UserWarning: Expected 2 arguments for function <function bad>, received 1' plus 'Expected at least 2 arguments … received 1' — and the app still launches, quietly filling the missing argument with None on every tick (reproduced in this run). One component, one argument.

$ import gradio as gr
import numpy as np
import time


def on_start():
    return "started at " + time.strftime("%H:%M:%S")


def keep(audio):                     # runs once, on the button click
    if audio is None:
        return None, "nothing recorded"
    sr, data = audio
    return (sr, data), "%d frames @ %d Hz = %.2f s" % (data.shape[0], sr, data.shape[0] / sr)


with gr.Blocks(title="Voice memo") as demo:
    gr.Markdown("## Voice memo — record, inspect, play back")
    with gr.Row():
        mic = gr.Microphone(label="memo")
        play = gr.Audio(label="playback", type="numpy", editable=False)
    info = gr.Textbox(label="clip info")
    stamps = gr.Textbox(label="recorder events")
    mic.start_recording(on_start, None, stamps, api_name="started")
    mic.stop_recording(lambda: "stopped at " + time.strftime("%H:%M:%S"), None, stamps, api_name="stopped")
    btn = gr.Button("Keep this memo", variant="primary")
    btn.click(keep, mic, [play, info], api_name="keep")

demo.launch(prevent_thread_lock=True, server_port=7923)
* Running on local URL:  http://127.0.0.1:7923
* To create a public link, set `share=True` in `launch()`.

Renders: a 'Voice memo' heading, a Row with the recorder labeled 'memo' next to a non-editable player labeled 'playback', then two Textboxes ('clip info', 'recorder events') and a primary Button 'Keep this memo'. Startup emits two UserWarnings naming the handler: 'Expected 1 arguments for function <function on_stop>, received 0' and 'Expected at least 1 arguments … received 0' — a recorder event wired with no input components still launches, it just cannot receive the clip.

The tuple returned by keep() lands in the player as a cache file, and because gr.Audio there has format=None the wav is passed straight through. Bonus, verified in the same run: a Microphone can also be an output — .click(push, None, mic) served the synthesized clip as '…/gradio/43a95c34…/audio.wav'. It is the format='wav' default doing that re-encode, not a rule about inputs.

Common flags

sources
Accepted by the signature, forwarded to super(), then overwritten by a hard-coded sources = ["microphone"] inside __init__. Verified: gr.Microphone(sources=["upload"]).sources reads ['microphone'] — no exception, not even a warning. Leave it out of your code.
format="wav"
The default here is "wav", not null like gr.Audio. It converts user audio on the way in when type="filepath" (an mp3 upload arrived at the handler as a 48 kB wav) and sets the container of anything you return. format=None opts out.
type="numpy" | "filepath"
"numpy" (default) hands your handler a tuple of (int sample rate, int16 array shaped (samples,) mono or (samples, channels)); "filepath" hands it a str pointing into the Gradio cache.
recording=False
Documented as: if True, the component is set to record audio from the microphone when the app loads. Verified to survive the constructor and reach the config as recording=True.
streaming=True
Requires 'microphone' in sources — trivially true here, so it cannot fail on a Microphone. The handler gets exactly one argument per chunk, and .stream() is registered with connection='stream' and stream_every defaulting to 0.5 s.
waveform_options=gr.WaveformOptions(...)
Recorder visuals: waveform_color, waveform_progress_color, trim_region_color, show_recording_waveform=False to hide the live trace, skip_length for the skip buttons, and sample_rate — 44100 by default — the rate edited audio is resampled to.
start_recording / pause_recording / stop_recording
The events a plain upload widget does not have. Each fires once per take rather than per chunk, which makes them the right place for transcription, scoring, timestamps or a save step.

History

2019 — a component named Microphone that did not give you audio

The first audio code in Gradio is commit f2814e0e ('updated preprocessing for images and added preprocessing for audio', 2019-06-22), shipped in the 0.7.8 sdist uploaded that day. It was class Microphone in gradio/inputs.py — and its preprocess ran generate_mfcc_features_from_audio_file, handing your model MFCC coefficients instead of a waveform. Audio recording in Gradio started life as a feature extractor.

2022 — Microphone becomes a template, not a class

gr.Audio was added to inputs.py and outputs.py in 1.1.0 (2020-08-10) while Microphone stayed a legacy class beside it. The 3.0.0 sdist (2022-05-16) deleted it from inputs.py and introduced gradio/templates.py, where `class Microphone(components.Audio)` now sets sources=["microphone"] and forwards 26 parameters to super() — a file of one-line components. class Mic, which passed the older singular source="microphone", was itself replaced by the alias `Mic = Microphone` in 3.12.0 (2022-11-29).

Fun facts

Pros & cons

pros

  • + One line gives a record-first UI with no upload tab to explain away — the value shape is always the same (sample_rate, array) tuple or one cache path
  • + It is the full gr.Audio class underneath: waveform_options, subtitles, buttons, playback_position, output mode, streaming and preprocess all behave identically
  • + The three recorder lifecycle events (start_recording, pause_recording, stop_recording) give you per-take hooks that .stream() cannot — timestamp the beginning, run work on the finished clip

cons

  • − sources is hard-coded and overriding it is silently ignored, so there is no configurable fallback for users who cannot or will not record
  • − format defaults to "wav" instead of Audio's None, so using one as an output re-encodes every clip you return — lossless, but a 1.5 s mp3 arrived as a 48 kB wav in this run
  • − It exists only for convenience: templates.py is a file of one-line wrappers, and nothing here tells you which behaviours are its own (the wav default) and which are inherited (everything else)

Takeaways

  1. 1Choose the component by input path: gr.Microphone() for capture only, gr.Audio(sources=["upload"]) for a dropzone, gr.Audio(sources=["upload", "microphone"]) for both — and never pass sources= to Microphone.
  2. 2One component, one handler argument. For the default type='numpy' that argument is the (sample_rate, ndarray) tuple, not its two halves.
  3. 3Wire .stop_recording() for per-take work and .stream() for live meters — the first fires once per recording, the second roughly twice a second.
  4. 4Remember the wav default: a Microphone used as an output converts what you return, and its format=None is how you stop it.
  5. 5Recorder events fire without the clip if you give them no input components, so pass mic into start/stop_recording handlers whenever the timestamp needs to match the take.

Related commands

← all learning