What you can make
Music and sound
Five tools make sound, two get it onto a clip, and this is the newest and roughest corner of Plnty.
If a board review needs something under it, **sonilo-v1.1** writes a track from a description. Give it a line like "warm lo-fi piano, slow tempo, vinyl dust" and it returns 10 seconds to 3 minutes of it, instrumental, licensed for commercial use, priced by the second. **elevenlabs-music-v2.5** and **elevenlabs-music-v2** go further than a bed. They write a full arrangement, with instruments that sound played rather than programmed, and with singing when the prompt asks for it, from R\&B to rock. The request is the same shape as sonilo's, a description and a length, plus an **Instrumental** switch for when you want no voice on it at all. Their price works differently and it is worth knowing before you run one. They are billed **per whole minute, rounded up**, and the model returns a little more audio than you ask for. A 60 second track is therefore billed as two minutes and costs 180 credits, where sonilo stays at 12 for 30 seconds. Each step of the length slider adds a minute. v2.5 is the newer of the two and they cost the same, so reach for v2 only if you prefer what it does. **elevenlabs-sfx** does the same for single sounds. Describe one, get a clean mp3 between 1 and 22 seconds, and switch on Loop for a version that repeats without an audible join, made for ambience beds. Fire it and walk away: generation survives a refresh and the audio card appears on the canvas when it is ready. **ace-instrument** is the experimental one. It drives ACE-Step through a set of semantic pills, and its gear panel exposes the model's real synthesis knobs, guidance and granularity and scheduler among them. Instrumental only, 8 seconds to 3 minutes. ## Getting it onto a clip Two routes, and they solve different problems. **mmaudio-v2** watches a silent video and generates audio synced to what happens on screen. What comes back is your same clip with the track baked in, as one playable mp4, anywhere from 1 to 30 seconds. Leave the prompt blank and it scores the clip on its own; write a direction and it follows. **audio-video-merge** is a small edit suite you drive yourself. Drop a canvas video, drag audio cards onto timeline lanes, trim and slide them, pull the edge handles for fades, set per-track gain, export an MP4. It runs ffmpeg compiled into your browser, so when the source is already MP4 the video stream is copied rather than re-encoded. Nothing degrades and nothing waits on a provider. ## What is missing There is no speech. No text-to-speech, no voiceover, no dialogue. The two ElevenLabs Music models do sing, but only as part of a song they are writing: you cannot hand them a line and get it read back. The generated audio is also the least developed part of Plnty. **ace-instrument** is openly experimental, and **mmaudio-v2** frequently returns something that is synchronised without being good. Both are being worked on. Use them for a scratch track and reach for a real composer when the piece matters.If a board review needs something under it, sonilo-v1.1 writes a track from a description. Give it a line like “warm lo-fi piano, slow tempo, vinyl dust” and it returns 10 seconds to 3 minutes of it, instrumental, licensed for commercial use, priced by the second.
elevenlabs-music-v2.5 and elevenlabs-music-v2 go further than a bed. They write a full arrangement, with instruments that sound played rather than programmed, and with singing when the prompt asks for it, from R&B to rock. The request is the same shape as sonilo’s, a description and a length, plus an Instrumental switch for when you want no voice on it at all.
Their price works differently and it is worth knowing before you run one. They are billed per whole minute, rounded up, and the model returns a little more audio than you ask for. A 60 second track is therefore billed as two minutes and costs 180 credits, where sonilo stays at 12 for 30 seconds. Each step of the length slider adds a minute. v2.5 is the newer of the two and they cost the same, so reach for v2 only if you prefer what it does.
elevenlabs-sfx does the same for single sounds. Describe one, get a clean mp3 between 1 and 22 seconds, and switch on Loop for a version that repeats without an audible join, made for ambience beds. Fire it and walk away: generation survives a refresh and the audio card appears on the canvas when it is ready.
ace-instrument is the experimental one. It drives ACE-Step through a set of semantic pills, and its gear panel exposes the model’s real synthesis knobs, guidance and granularity and scheduler among them. Instrumental only, 8 seconds to 3 minutes.
Getting it onto a clip
Section titled “Getting it onto a clip”Two routes, and they solve different problems.
mmaudio-v2 watches a silent video and generates audio synced to what happens on screen. What comes back is your same clip with the track baked in, as one playable mp4, anywhere from 1 to 30 seconds. Leave the prompt blank and it scores the clip on its own; write a direction and it follows.
audio-video-merge is a small edit suite you drive yourself. Drop a canvas video, drag audio cards onto timeline lanes, trim and slide them, pull the edge handles for fades, set per-track gain, export an MP4. It runs ffmpeg compiled into your browser, so when the source is already MP4 the video stream is copied rather than re-encoded. Nothing degrades and nothing waits on a provider.
What is missing
Section titled “What is missing”There is no speech. No text-to-speech, no voiceover, no dialogue. The two ElevenLabs Music models do sing, but only as part of a song they are writing: you cannot hand them a line and get it read back.
The generated audio is also the least developed part of Plnty. ace-instrument is openly experimental, and mmaudio-v2 frequently returns something that is synchronised without being good. Both are being worked on. Use them for a scratch track and reach for a real composer when the piece matters.