๐ŸŽ™๏ธ Audio & Speech

Speech-to-text, text-to-speech, and audio understanding. 40 models.

G

Google: Gemini 2.5 Flash Preview TTS

Google

A specialized preview of Gemini 2.5 Flash with text-to-speech capabilities. Combines fast inference with high-quality voice synthesis for real-time audio applications.

Streaming
G

Google: Gemini 2.5 Pro Preview TTS

Google

A specialized preview of Gemini 2.5 Pro with text-to-speech capabilities. Designed for applications requiring high-quality voice synthesis alongside advanced language understanding.

Streaming
G

Google: Gemini 3.1 Flash TTS Preview

Google

Gemini 3.1 Flash TTS Preview is a capable language model available on LangMart via Google for general-purpose text generation and analysis tasks.

Streaming
G

Google: Gemini 3.5 Transcribe

Google

Gemini 3.5 Transcribe is a capable language model available on LangMart via Google for general-purpose text generation and analysis tasks.

Streaming
G

Google: Gemini Audio Understanding

Google

Audio analysis model.

Vision
G

Groq: Distil Whisper Large V3 EN

Groq

Groq: Distil Whisper Large V3 EN is a language model provided by the provider. This model offers advanced capabilities for natural language processing tasks.

Streaming Vision
G

Groq: Orpheus Arabic Saudi

Groq

A specialized Arabic language model from Canopy Labs focused on Saudi Arabian dialect and culture. Optimized for Saudi Arabic text generation and understanding.

G

Groq: Orpheus V1 English

Groq

Canopy Labs' Orpheus V1 model for English language tasks. A lightweight model optimized for fast inference on Groq infrastructure.

G

Groq: PlayAI TTS

Groq

PlayAI's text-to-speech model hosted on Groq for ultra-fast audio generation. Provides high-quality voice synthesis with low latency.

G

Groq: PlayAI TTS Arabic

Groq

PlayAI's Arabic-specialized text-to-speech model hosted on Groq. Optimized for high-quality Arabic voice synthesis.

G

Groq: Whisper Large V3

Groq

OpenAI's Whisper Large V3 speech-to-text model hosted on Groq for ultra-fast transcription. Supports multiple languages and audio formats.

G

Groq: Whisper Large V3 Turbo

Groq

Groq: Whisper Large V3 Turbo is a language model provided by the provider. This model offers advanced capabilities for natural language processing tasks.

Streaming Vision
G

Groq: Whisper V3 Large

Groq

Whisper V3 Large is a audio processing model for speech synthesis and audio understanding. Developed by Groq, this model is optimized for its specific use case category.

Streaming Vision
M

Mistral AI: Voxtral Mini

Mistral AI

Voxtral Mini is a audio processing model for speech synthesis and audio understanding. Developed by Mistral AI, this model is optimized for its specific use case category.

Streaming Vision
M

Mistral AI: Voxtral Small

Mistral AI

Voxtral Small is a audio processing model for speech synthesis and audio understanding. Developed by Mistral AI, this model is optimized for its specific use case category.

Streaming Vision
O

OpenAI: Gpt 4O Audio Preview 2024 12 17

OpenAI

OpenAI model: gpt-4o-audio-preview-2024-12-17 This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solv...

Streaming
O

OpenAI: Gpt 4O Audio Preview 2025 06 03

OpenAI

OpenAI model: gpt-4o-audio-preview-2025-06-03 This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solv...

Streaming
O

OpenAI: Gpt 4O Mini Audio Preview

OpenAI

OpenAI model: gpt-4o-mini-audio-preview This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving ta...

Streaming
O

OpenAI: Gpt 4O Mini Audio Preview 2024 12 17

OpenAI

OpenAI model: gpt-4o-mini-audio-preview-2024-12-17 This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem...

Streaming
O

OpenAI: Gpt 4O Mini Transcribe

OpenAI

OpenAI model: gpt-4o-mini-transcribe This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving tasks...

Streaming
O

OpenAI: Gpt 4O Mini Transcribe 2025 03 20

OpenAI

OpenAI model: gpt-4o-mini-transcribe-2025-03-20 This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-so...

Streaming
O

OpenAI: Gpt 4O Mini Transcribe 2025 12 15

OpenAI

OpenAI model: gpt-4o-mini-transcribe-2025-12-15 This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-so...

Streaming
O

OpenAI: Gpt 4O Transcribe

OpenAI

OpenAI model: gpt-4o-transcribe This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving tasks.

Streaming
O

OpenAI: Gpt 4O Transcribe Diarize

OpenAI

OpenAI model: gpt-4o-transcribe-diarize This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving ta...

Streaming
O

OpenAI: Gpt Audio

OpenAI

OpenAI model: gpt-audio This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving tasks.

Streaming
O

OpenAI: Gpt Audio 2025 08 28

OpenAI

OpenAI model: gpt-audio-2025-08-28 This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving tasks.

Streaming
O

OpenAI: Gpt Audio Mini

OpenAI

OpenAI model: gpt-audio-mini This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving tasks.

Streaming
O

OpenAI: Gpt Audio Mini 2025 10 06

OpenAI

OpenAI model: gpt-audio-mini-2025-10-06 This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving ta...

Streaming
O

OpenAI: Gpt Audio Mini 2025 12 15

OpenAI

OpenAI model: gpt-audio-mini-2025-12-15 This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-solving ta...

Streaming
O

OpenAI: GPT-4o Audio Preview

OpenAI

Preview of GPT-4o with audio processing capabilities. This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex prob...

Vision Tools Streaming Reasoning
O

OpenAI: GPT-4o Audio Preview (October 2024)

OpenAI

October release of audio-enabled GPT-4o preview. This model supports multimodal capabilities including vision and image understanding. It features advanced reasoning capabilities for complex problem-s...

Vision Tools Streaming Reasoning
O

OpenAI: GPT-4o Mini TTS

OpenAI

GPT-4o Mini TTS is a audio processing model for speech synthesis and audio understanding. Developed by OpenAI, this model is optimized for its specific use case category.

Streaming Vision
O

OpenAI: TTS

OpenAI

TTS is a audio processing model for speech synthesis and audio understanding. Developed by OpenAI, this model is optimized for its specific use case category.

Streaming Vision
O

OpenAI: TTS HD

OpenAI

TTS HD is a audio processing model for speech synthesis and audio understanding. Developed by OpenAI, this model is optimized for its specific use case category.

Streaming Vision
O

OpenAI: Whisper

OpenAI

Whisper is a audio processing model for speech synthesis and audio understanding. Developed by OpenAI, this model is optimized for its specific use case category.

Streaming Vision
O

LangMart: Mistral: Voxtral Small 24B 2507

Openrouter

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translati...

O

LangMart: OpenAI: GPT-4o Audio

Openrouter

OpenAI: GPT-4o Audio is a conversational AI model designed for multi-turn dialogue and interactive tasks. Developed by LangMart, this model is optimized for its specific use case category.

Streaming Vision
O

OpenAI: GPT Audio

Openrouter

OpenAI: GPT Audio is a capable language model available on LangMart via OpenRouter for general-purpose text generation and analysis tasks.

Streaming
O

OpenAI: GPT Audio Mini

Openrouter

OpenAI: GPT Audio Mini is a capable language model available on LangMart via OpenRouter for general-purpose text generation and analysis tasks.

Streaming
T

Together AI: Whisper Large v3

Together AI

Whisper Large v3 is a conversational AI model designed for multi-turn dialogue and interactive tasks. Developed by Together AI, this model offers reliable performance for diverse applications.

Streaming Vision