Model reference · open weights

acestep-captioner

Available as managed deployment Audio ACE-Step Music / audio 1 variants 1k dl/mo

acestep-captioner is an open-weight audio or speech model from ACE-Step. AxForge deploys and operates it for you on dedicated EU-owned hardware — with the licence handled where one is required.

Available as managed deployment — configured and operated for you on dedicated EU hardware, quoted per deployment.

What it is

MakerACE-Step
TypeAudio & music
TaskMusic / audio
Parameters (lead)10.7B
Runs withtransformers
Released2026-01-23
Popularity1k downloads / month
LicenceOpen weights

About

What acestep-captioner is

Description

ACE-Step Captioner is the annotation model used by ACE-Step v1.5 for training data labeling. It is a professional-grade music captioning model that generates detailed, structured descriptions of audio content.

Performance

🏆 Accuracy surpasses Gemini Pro 2.5 in music description tasks

Key Features

  • 🎼 Musical Style Analysis - Identifies genres, sub-genres, and stylistic influences
  • 🎸 Instrument Recognition - Detects and describes 1000+ instrument types and combinations
  • 🎭 Structure & Progression - Analyzes musical arrangement including intro, verse, chorus, bridge, climax, and outro
  • 🔊 Timbre Description - Captures tonal qualities, textures, and sonic characteristics
  • 📝 Rich Vocabulary - Supports 1000+ descriptive terms for comprehensive music annotation

Usage

The usage is the same as Qwen2.5 Omni-7B.

Prompt Format

Use the following prompt to caption audio:

*Task* Describe this audio in detail

Output Format

The model generates natural language descriptions covering multiple aspects of the music.

Example Output

A melancholic indie folk track featuring fingerpicked acoustic guitar
as the primary instrument. The song opens with a sparse, contemplative
intro before the vocals enter with a breathy, intimate delivery.
The arrangement gradually builds through the verse, adding subtle
string pads and a gentle kick drum. The chorus lifts with layered
harmonies and a warmer, fuller texture. The bridge introduces a
key change and emotional climax before returning to the stripped-down
acoustic arrangement for the outro.

Descriptive Capabilities

Musical Styles (Examples)

CategoryStyles
ElectronicAmbient, Techno, House, Drum & Bass, Synthwave, IDM, Downtempo
RockAlternative, Indie, Post-Rock, Progressive, Psychedelic, Grunge
PopSynth-pop, Electropop, Dream Pop, Art Pop, Indie Pop
ClassicalOrchestral, Chamber, Minimalist, Neo-Classical, Cinematic
WorldLatin, African, Middle Eastern, Asian Traditional, Celtic
JazzFusion, Smooth, Bebop, Modal, Free Jazz
Hip-HopTrap, Boom Bap, Lo-fi, Instrumental, Cloud Rap

Instruments (1000+ Supported)

CategoryExamples
StringsAcoustic Guitar, Electric Guitar, Violin, Cello, Bass, Harp, Mandolin
KeysPiano, Synthesizer, Organ, Rhodes, Wurlitzer, Mellotron
PercussionDrums, Electronic Drums, Congas, Bongos, Timpani, Vibraphone
WindSaxophone, Trumpet, Flute, Clarinet, Oboe, French Horn
ElectronicSynth Bass, Pad, Lead, Arpeggiator, Sampler, 808, 303

Structure Analysis

  • Intro / Outro - Opening and closing sections
  • Verse / Pre-Chorus / Chorus - Main song structure
  • Bridge / Break - Transitional sections
  • Build-up / Drop / Climax - Dynamic progression
  • Interlude / Solo - Instrumental passages

Timbre Descriptions

DimensionDescriptors
TextureWarm, Bright, Dark, Crisp, Muddy, Clean, Distorted, Saturated
SpaceReverberant, Dry, Spacious, Intimate, Cavernous, Tight
DynamicsPunchy, Soft, Aggressive, Gentle, Compressed, Dynamic
CharacterEthereal, Gritty, Smooth, Raw, Polished, Organic, Synthetic

Use Cases

  • Music AI Training - Generate high-quality captions for music generation models
  • Music Information Retrieval - Create searchable metadata for audio databases
  • Content Moderation - Analyze and categorize music content
  • Music Education - Provide detailed analysis for learning purposes
  • Audio Production - Document and describe sound design elements

From the published model card. Full card on the HuggingFace links in the sidebar.

Using it via the API

Call it like any OpenAI endpoint

Once AxForge deploys acestep-captioner for you, it answers on the OpenAI-compatible API — the same base URL and keys as every other model. (acestep-captioner below is illustrative; you get the exact model name on deployment.)

$ curl -sS https://api.axforge.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $AXFORGE_API_KEY" \
  -F model="acestep-captioner" -F file=@audio.mp3

Create an account — your API key is available in the console. 5M tokens/month currently included with every new account at launch.

© 2026 AxForge · EU-hosted AI infrastructure Pricing Docs Trust Privacy Terms