Stop gluing cv2 and Flask together to accept an image — gr.Image hands your function pixels, with three type modes and a one-line webcam.
The most-forked Gradio demos on the Hub all begin the same way: someone wrapped a model in four lines, and the first line was gr.Image().
gr.Image() is both an input and an output component. As an input it renders a drop area with up to three sources — upload, webcam, clipboard — and converts whatever arrives into the shape your function asked for: type="numpy" (default) gives an RGB uint8 array of shape (h, w, 3), type="pil" a PIL.Image, and type="filepath" a plain path on disk. As an output, set interactive=False and hand it back a numpy array, a PIL image, a path, or even a URL — Gradio saves it to the cache and serves it to the browser. Six event listeners are wired in: .change(), .input(), .clear(), .select(), .upload() and .stream() for live webcam frames.
Every computer-vision demo — classifiers, detection, segmentation, upscalers, inpainting — needs the same handshake: get pixels from a human, convert them, show a result. Getting this right by hand costs a multipart-form file handler, MIME sniffing, a PIL/np roundtrip, and a temp-file policy. gr.Image() is that handshake, already written, tested, and wired to a queue that survives concurrent users. The two parameters that matter are sources (which UI affordances appear) and type (what your function receives); everything else — image_mode, format, webcam mirroring, watermarks — is tuning. A classifier that took an afternoon with Flask and a <form> tag is eleven lines of Python here.
import gradio as gr
with gr.Blocks(title='classifier') as demo:
gr.Markdown("## Plant leaf disease classifier")
with gr.Row():
with gr.Column():
img_in = gr.Image(label='Leaf photo',
sources=['upload', 'webcam', 'clipboard'],
type='numpy', image_mode='RGB')
with gr.Column():
conf = gr.Slider(0, 1, 0.5, label='Min confidence')
btn = gr.Button('Classify', variant='primary')
lbl = gr.Label(label='Prediction', num_top_classes=3)
md = gr.Markdown('')
def classify(image, min_conf):
h, w = image.shape[:2]
scores = {'healthy': 0.84, 'leaf_rust': 0.11, 'powdery_mildew': 0.05}
top = {k: v for k, v in scores.items() if v >= min_conf}
return top, f'input tensor: {h}×{w}×{image.shape[2]} ({image.dtype})'
btn.click(classify, [img_in, conf], [lbl, md])
demo.launch()* Running on local URL: http://127.0.0.1:7880
* To create a public link, set `share=True` in `launch()`.
Rendered UI: a header, then one Row with two Columns — left: an image drop area labeled "Leaf photo" with upload / webcam / clipboard tabs; right: a labeled Slider "Min confidence" and a green primary Button "Classify". Below the Row: a Label panel and an empty Markdown strip. The function receives a real numpy array: (verified by calling it with a 384×256×3 uint8 array) with min_conf=0.05 it returns {'healthy': 0.84, 'leaf_rust': 0.11, 'powdery_mildew': 0.05} and the note "input tensor: 384×256×3 (uint8)"; at min_conf=0.20 only {'healthy': 0.84} survives the threshold.
Every image demo is this shape: Image in, some control widgets, Label/Markdown/Image out. Swap the dict in scores for your model's softmax and it's a real product.
import gradio as gr
with gr.Blocks(title='resizer') as demo:
gr.Markdown("## Thumbnailer")
img = gr.Image(label='Image', type='pil', format='jpeg', sources=['upload'])
size = gr.Radio([64, 128, 256], label='Thumbnail size', value=128)
out_img = gr.Image(label='Thumbnail', interactive=False)
meta = gr.Textbox(label='Metadata', lines=1)
def thumb(im, s):
small = im.resize((s, s))
return small, f'original {im.width}×{im.height} → {s}×{s} (PIL {im.mode})'
img.upload(thumb, [img, size], [out_img, meta])
demo.launch()* Running on local URL: http://127.0.0.1:7881 * To create a public link, set `share=True` in `launch()`. Rendered UI: an upload-only image drop area labeled "Image", a 3-option Radio "Thumbnail size", then a second (non-interactive) Image labeled "Thumbnail" and a single-line Textbox "Metadata". With type="pil" your function gets a PIL.Image (verified: a 256×384 RGB source resized to 128×128, mode RGB) — no numpy conversion needed. format="jpeg" controls how numpy/PIL results are re-encoded when sent back to the browser.
.upload() fires only after a file lands (4.4.0 fixed it not firing at all — PR #6441); use .change() when you also want clicks on the preview to retrigger.
import gradio as gr
with gr.Blocks(title='stream') as demo:
cam = gr.Image(label='Camera', sources=['webcam'], streaming=True)
info = gr.Textbox(label='Frame info')
cam.stream(lambda f: (f'frame {f.shape[0]}×{f.shape[1]}'
if f is not None else 'none'),
cam, info)
demo.launch()* Running on local URL: http://127.0.0.1:7882 * To create a public link, set `share=True` in `launch()`. Rendered UI: a webcam-only Image card labeled "Camera" (browser asks for camera permission on first use, every captured frame mirrors by default) and a Textbox "Frame info" that updates per streamed frame. Caught by construction, not by a browser: gr.Image(streaming=True, sources=['upload', 'webcam']) refuses to build with ValueError: Image streaming only available if sources is ['webcam']. Streaming not supported with multiple sources. That guard (check_streamable) runs when the Blocks close — you find out at startup, not in front of a user.
streaming=True also changes the resolved default: with streaming on, sources defaults to ['webcam'] only, and the output side of a streaming pipeline re-encodes frames as base64.
| Flag | Meaning |
|---|---|
type | "numpy" (default) → RGB uint8 (h, w, 3); "pil" → PIL.Image; "filepath" → str path. Swap per demo, zero code changes. |
sources | Subset of ['upload', 'webcam', 'clipboard']; default is all three (or exactly ['webcam'] when streaming=True). |
image_mode | PIL mode used to load pixels: "RGB" default, "L" for grayscale, "RGBA" to keep alpha; None infers from the file. |
format | Re-encode target for numpy/PIL results ("png", "jpeg", "webp"…); default webp. Only affects server→browser encoding. |
streaming | Live webcam mode: your function runs per frame via .stream(); output Images stream base64. Valid only with sources=['webcam']. |
webcam_options | WebcamOptions(mirror=True, constraints={…}) — flip the mirror or pin a camera resolution/dimensions via MediaTrack constraints. |
watermark | WatermarkOptions image stamped bottom-right of displayed values (5.45.0+, PR #11831). Display-only; not baked into files you return. |
gr.Image predates the whole Blocks era, but Gradio 3.4 (Sept 2022) made it a star: the same release added type-ahead Sliders and the Sketchpad/Paint tools for inpainting. The tool= parameter era ended fast — Gradio 4.0 (Oct 2023) removed tool= from the signature, and 4.5.0 (Nov 2023, PR #6169) reintroduced that functionality as the dedicated gr.ImageEditor component with brush/eraser dataclasses.
Early Gradio handed images to your function in whatever shape the frontend had, and every demo started with PIL.Image.open(...). The type= parameter moved the conversion into the component's preprocess — your function now states its contract (numpy, PIL, or filepath) and Gradio guarantees it. The 2024+ cleanup finished the job: source → sources (lists), and tool= replaced by ImageEditor.
The browser never sends your function a file straight away. upload → the frontend POSTs multipart bytes → routes.py wraps them in an ImageData (a FileData with base64 image data) → preprocess() converts to your type= target — for numpy, image_utils.preprocess_image applies image_mode first (so "L" yields (h, w) uint8, "RGB" (h, w, 3)). On the way back, postprocess() takes your numpy/PIL/path value, writes it to the cache dir in format (webp default; a path with a valid extension is passed through), and returns a FileData that points at the cached file. The six event listeners are just typed wrappers around the same payload — .select() even carries gr.SelectData with the clicked pixel coordinates, so you can build a click-to-zoom demo without any JavaScript.