Turn English speech into punctuated text with word timestamps using an NVIDIA open ASR model.
Audio & Voice

Use text and optional audio to explore unusual sound generation and transformation.
Best for: Musicians, sound designers, researchers, and curious creators exploring controllable generative audio.
Turn English speech into punctuated text with word timestamps using an NVIDIA open ASR model.
Transform a live sound into a new playable texture while keeping its musical shape.
Run multilingual speech recognition or translation with an open NVIDIA model.