Model details
Active Speaker Detection
Nvidia's Active Speaker Detection is published through the NVIDIA NGC catalog under the Maxine SDK collection, placing it alongside real-time conferencing and avatar-streaming assets such as DeepStream. The URL path within NGC groups the offering with Maxine, suggesting an intended role in live video pipelines that need to determine who in a frame is currently talking. Because the available excerpt only captures NGC navigation chrome, concrete architectural choices, training corpus, and benchmark numbers cannot be cited as verified model facts.
Adjacent academic work on deep audio-visual attention demonstrates the broader research direction this product fits into: jointly reasoning over face crops and their paired audio to predict speaking activity. This framing helps explain why such a model is valuable for video conferencing, broadcast production, and accessibility tooling, where reliable per-frame speaker labels enrich downstream effects, transcription, or analytics. Without a model card in the provided evidence, claims about parameter counts, context handling, supported resolutions, or quantitative leaderboard positions are deliberately omitted.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- nvidia/active-speaker-detection
- Release date
- Apr 16, 2026
- Last updated
- Apr 16, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 4,096 tokens
- Context window
- 0 tokens