gr.Audio — record, upload, edit and play audio in one component

One component, three jobs — and a (sample_rate, array) contract you can actually rely on.

Your function receives the audio in a different shape depending on one constructor argument, and the default is not the one most tutorials show.

What it does

gr.Audio is a player, a recorder and a file dropzone in one component. As an input it hands your function either a (sample_rate, np.int16 array) tuple — type='numpy', the default — or a str filepath (type='filepath'). As an output it accepts a path, a URL, a bytes object, or that same tuple, and renders an HTML5 player; tuple values are written to a .wav in the Gradio cache before the browser sees them. sources decides what the user gets: 'upload', 'microphone', or both (the default). The waveform is editable by default, so users can trim before the value ever reaches you.

Why it matters

Voice and music demos live or die on upload friction, and gr.Audio is the component behind the 'record your voice' Spaces — the Whisper demos, TTS comparisons and speaker-ID toys you have used. It absorbs the parts that normally eat an afternoon: base64 vs multipart, browser codec support, temp-file lifetimes, and the difference between a live recording and an uploaded file. Get type= right and your handler is ten lines of numpy away from any model.

Examples

import gradio as gr
import numpy as np


def inspect_audio(audio):            # default type: audio = (sample_rate, np.int16 array)
    if audio is None:
        return {"status": "no clip"}
    sr, data = audio
    mono = data.mean(axis=1) if data.ndim > 1 else data
    rms = float(np.sqrt((mono.astype(np.float64) ** 2).mean()))
    return {
        "sample_rate": sr,
        "samples": int(data.shape[0]),
        "channels": int(data.shape[1]) if data.ndim > 1 else 1,
        "duration_s": round(data.shape[0] / sr, 3),
        "peak_dbfs": round(20 * np.log10(max(abs(mono).max(), 1) / 32768), 1),
        "rms_dbfs": round(20 * np.log10(max(rms, 1) / 32768), 1),
    }


with gr.Blocks(title="Speech clip QA") as demo:
    gr.Markdown("## Voice dataset QA — record or drop a clip")
    with gr.Row():
        clip = gr.Audio(label="clip", sources=["upload", "microphone"], type="numpy")
        report = gr.JSON(label="signal report")
    clip.change(inspect_audio, clip, report, api_name="inspect")

demo.launch(prevent_thread_lock=True, server_port=7911)
* Running on local URL:  http://127.0.0.1:7911
* To create a public link, set `share=True` in `launch()`.

Renders: a 'Voice dataset QA' header, then a Row holding an audio widget labeled 'clip' (drop-a-file box plus a Microphone tab, because both sources are enabled) and the JSON panel labeled 'signal report'. Driving the live /inspect endpoint with gradio_client and a 2 s 440 Hz test WAV returned exactly {'sample_rate': 16000, 'samples': 32000, 'channels': 1, 'duration_s': 2.0, 'peak_dbfs': -8.7, 'rms_dbfs': -11.7}.

Verified on gradio 6.27.0 — that dict is the real event payload, not an illustration. type='numpy' is the constructor default, so this snippet is what you get if you never set type at all.

import gradio as gr


def normalize(clip_path):            # type='filepath': your handler gets one plain str
    return clip_path, clip_path


with gr.Blocks() as demo:
    src = gr.Audio(label="drop any container (.wav, .flac, .ogg)",
                   sources=["upload"], type="filepath", format="mp3")
    path = gr.Textbox(label="path your handler receives")
    play = gr.Audio(label="the mp3 Gradio produced")
    src.change(normalize, src, [path, play], api_name="normalize")

demo.launch(prevent_thread_lock=True, server_port=7912)
* Running on local URL:  http://127.0.0.1:7912
* To create a public link, set `share=True` in `launch()`.

Uploading the same 2 s WAV through the live event returned ('…/gradio/79b4de57…/tone440.mp3', '…/gradio/cea082cb…/tone440.mp3') — same stem, swapped extension, 6 489 bytes against the 64 044-byte WAV input. UI: an upload dropzone labeled with the container list, a Textbox echoing the converted path your function received, and a second player labeled 'the mp3 Gradio produced'.

type='filepath' returns the path inside the Gradio cache, not the user's original file. Note what is NOT true here: format= only re-encodes on the way in when the uploaded extension differs from 'mp3' — a .flac upload with format='flac' would be passed through untouched.

import gradio as gr
import numpy as np


def chime(seconds):                  # generate audio: return (sample_rate, np.int16 array)
    sr = 44100
    t = np.linspace(0, seconds, int(sr * seconds), endpoint=False)
    envelope = np.exp(-3 * t)        # decaying 440 Hz bell
    pcm = (0.6 * np.sin(2 * np.pi * 440 * t) * envelope * 32767).astype(np.int16)
    return (sr, pcm)


with gr.Blocks() as demo:
    dur = gr.Slider(0.2, 3.0, value=1.2, step=0.1, label="seconds")
    btn = gr.Button("Synthesize", variant="primary")
    out = gr.Audio(label="generated chime", autoplay=True, editable=False,
                   waveform_options=gr.WaveformOptions(waveform_color="#009245",
                                                       waveform_progress_color="#111827"))
    btn.click(chime, dur, out, api_name="chime")

demo.launch(prevent_thread_lock=True, server_port=7913)
* Running on local URL:  http://127.0.0.1:7913
* To create a public link, set `share=True` in `launch()`.

A /chime call with seconds=1.2 returned '…/gradio/470145ad…/audio.wav'. Reading that file back with wave: 52 920 frames at 44 100 Hz = 1.2 s exactly, 105 884 bytes, mono. UI: a Slider labeled 'seconds' (0.2 to 3.0, default 1.2), a primary Button 'Synthesize', and a non-editable player labeled 'generated chime' whose waveform is drawn in the two colors passed to gr.WaveformOptions.

A tuple return is the universal output format: postprocess writes it via save_audio_to_cache using format or 'wav'. The browser will ignore autoplay=True until the user has interacted with the page at least once.

import gradio as gr
import numpy as np


def live_level(audio):               # ONE argument — the (sample_rate, ndarray) tuple
    if audio is None:
        return "no signal"
    sr, chunk = audio
    if chunk is None or len(chunk) == 0:
        return "silence"
    rms = float(np.sqrt((chunk.astype(np.float64) ** 2).mean()))
    return f"{20 * np.log10(max(rms, 1) / 32768):.1f} dBFS"


with gr.Blocks() as demo:
    mic = gr.Audio(sources=["microphone"], streaming=True, type="numpy", label="live mic")
    level = gr.Textbox(label="level")
    mic.stream(live_level, mic, level, api_name="level")

demo.launch(prevent_thread_lock=True, server_port=7914)
* Running on local URL:  http://127.0.0.1:7914
* To create a public link, set `share=True` in `launch()`.

The served /config confirms one Audio component (streaming=True, sources=['microphone']) plus a Textbox, with exactly one dependency: [('level', [[1, 'stream']])] — i.e. the stream event on component 1. UI: a mic-only audio widget labeled 'live mic' with a Record button, and a Textbox whose text updates on every streamed chunk.

The wrong signature is instructive: with def live_level(sr, chunk) gradio logs 'UserWarning: Expected 2 arguments for function <function live_level>, received 1' and then 'Unexpected argument. Filling with None.' — the textbox silently fills with None on every tick. One component, one argument.

Flags

FlagMeaning
type="numpy" | "filepath"The preprocess contract. 'numpy' (default) gives you (int sample_rate, np.int16 array of shape (samples,) or (samples, channels)); 'filepath' gives a str. Anything else raises ValueError at construction.
sourcesList of 'upload' / 'microphone'; None means ['upload', 'microphone'], or ['microphone'] when streaming=True. A bare string like sources='upload' is accepted and wrapped into a list (verified).
format="wav" | "mp3"The only two legal values — format='ogg' raises ValueError at construction (verified). None means no conversion; a numpy/bytes output with format=None lands as a .wav anyway.
streaming=TrueRequires 'microphone' in sources, otherwise ValueError. The handler gets ONE argument (the audio value); the frontend ships chunks to your .stream() event.
autoplay / loop / playback_positionOutput-side playback: autoplay is blocked by browsers until the page has been interacted with, loop replays from the end, playback_position both sets the start offset and reports the current one (added in 6.1.0).
waveform_options=gr.WaveformOptions(...)waveform_color, waveform_progress_color, trim_region_color, skip_length (percent jumped by the skip buttons), show_recording_waveform, and sample_rate — the rate the edited audio is resampled to, 44100 by default.
buttons=["download", "share"] | [] | [gr.Button(...)]Toolbar contents; defaults to download + share. buttons=[] removes both, and a gr.Button instance you pass appears in the toolbar and fires its own .click() handlers.
subtitles=path_or_dictsAn .srt/.vtt/.json path or URL, or a list of {'text', 'timestamp': [start, end]} dicts. Dict lists are normalized to {'start','end','text'} at construction (verified); a missing/invalid file raises ValueError immediately.

2019 — a mic input that did not give you audio

The first audio code in gradio is commit f2814e0e ('updated preprocessing for images and added preprocessing for audio', 2019-06-22), shipped in the 0.7.8 sdist uploaded that same day. It was a class called Microphone — and its preprocess did not return audio at all: it ran generate_mfcc_features_from_audio_file and handed your model MFCC coefficients. gradio 0.1.0 had been on PyPI since 2019-02-19.

2020-08 — gr.Audio arrives

Commit 919d01326 ('big changes', 2020-08-05) added class Audio to both gradio/inputs.py and gradio/outputs.py, plus static/js/interfaces/input/audio.js and the output audio CSS. It shipped in gradio 1.1.0 on 2020-08-10 — 1.0.5 and 1.0.6, both uploaded five days earlier, still only carried Microphone. The 2020 design had no sources= and no format=: a single source="upload" string, type in {'numpy','file','mfcc'}, and outputs that always became base64 wav.

preprocess → your function → postprocess, and the streaming encoder

preprocess is short and literal: type='numpy' calls processing_utils.audio_from_file and returns the (rate, int16 array) tuple; type='filepath' returns payload.path unchanged unless format= differs from the uploaded extension, in which case it decodes and re-writes a sibling file (a .wav uploaded to format='mp3' comes out tone440.mp3). postprocess branches on the four accepted types: bytes are saved to cache, a tuple goes through save_audio_to_cache (format or 'wav'), a str/Path either passes through, gets converted to your format=, or gets re-encoded to wav when ffmpeg says it is not browser-playable, and anything else raises ValueError: Cannot process <value> as Audio. Streaming output is a different machine entirely: on the first chunk Gradio registers an _EncoderSlot keyed by the output id, decodes each chunk to PCM, feeds a per-stream AAC encoder, and returns MediaStreamChunk segments whose duration is computed from the frame count (len(frames) * encoder.frame_duration) so the HLS playlist's #EXTINF stays honest. That is why streamed audio plays through the plain browser player rather than the waveform — and why an ended stream has to detach its encoder under a lock.

Fun facts

Pros

Cons

Takeaways