You can remove vocals from a song on Aihangsoft in one run: upload the track, press Separate vocals, and download two MP3 files, one with the isolated vocal and one with the instrumental. No install and no account are needed, but a guest gets a single free run a day, capped at 30 seconds of audio.
Two things decide whether this tool suits you. First, it is a Cloud AI tool, so your audio is uploaded to our server and forwarded to Volcengine AI MediaKit instead of being processed in the browser. Second, a free run is small and slow: 30 seconds of audio costs 5 credits, a free account gets 5 credits a day, and even a 5 second clip takes about 105 seconds because the service adds roughly 100 seconds of fixed overhead. The sections below cover what you get, what it costs, and the cases where a different tool is the better answer.
Most people arrive with one of two jobs: a backing track for karaoke, or a clean vocal for a remix, a sample or a transcript. Both are possible without installing anything. What is easy to miss is that two outputs is the whole story: no extra stems, no separation strength to tune, and no way to fit a full three minute song into the free allowance.
Vocal Remover
Splits a track into an acapella and an instrumental as two downloadable MP3 files. A Cloud AI tool, so your audio is uploaded for processing.
What you get: two MP3 tracks
Every successful run returns exactly two files, not a short preview or a link that vanishes with the page.
| Output | What is in the file | What people use it for |
|---|---|---|
| Vocals (acapella) | The vocal line, with the instruments largely removed | Remixes, sampling, lyric checks, transcription |
| Instrumental (backing track) | Drums, bass, guitars, keys and everything else | Karaoke, practice, cover recordings |
Each file has its own player and download button, so you can take one or both. Names come from your upload, so song.mp3 returns as song-vocals.mp3 and song-instrumental.mp3.
Results are temporary: 30 minutes for a guest, 2 hours for a free account and 24 hours on Pro. Save the files before you close the tab, because after that window there is no later re-download and the only way back is another run.
This is a cloud tool, so your audio is uploaded
Aihangsoft has 28 tools that run entirely in your browser: a file is read into page memory, processed there, and handed back as a download, with nothing transmitted. This is not one of them. Vocal separation needs a model too large to run in a browser tab, so your track is uploaded to our server, forwarded to Volcengine AI MediaKit, and the two results are written back for you to fetch.
| The 28 browser tools | This vocal remover | |
|---|---|---|
| Where your audio goes | Stays on your device | Uploaded to our server |
| What processes it | Your browser, using Canvas and WebAssembly | Volcengine AI MediaKit, a cloud AI service |
| How long results are kept | Not stored at all | 30 minutes as a guest, 2 hours free, 24 hours on Pro |
| Account | Never needed | Not required, but signing in raises the daily allowance and the length limit |
| Works offline | Yes, after the first load | No |
Treat that as a retention window rather than a promise about logs or backups. Do not upload an unreleased master, a private voice memo or a confidential interview, or anything you would not email to a stranger. A song you bought, a demo your own band recorded, or a clip you captured yourself is a reasonable trade.
How to remove vocals from a song in three steps
- Add your track. Drop in an audio or video file, one at a time, up to 30 MB. The duration is read in the browser first, so a clip longer than your plan allows is flagged and the button disabled before anything is uploaded.
- Press Separate vocals. There is nothing to configure, so there is no wrong setting to pick. The page uploads the file, submits the job, and then polls every 3 seconds while printing how long you have waited.
- Download both tracks. When the run completes, the vocals and the instrumental appear with their own players and download buttons. Save whichever you need before the retention window closes.
Nothing here is destructive and your original file is untouched. What you cannot do is change your mind halfway: there is no cancel button and no way to pause a running job.
What audio and video formats you can upload
The upload accepts both audio and video, and video is converted to audio on the way in rather than asking you to extract the sound yourself first.
| Input type | Formats | What happens |
|---|---|---|
| Audio | MP3, WAV, M4A, AAC, FLAC, OGG | Separated directly |
| Video | MP4, MOV, AVI, MKV, WebM, M4V, MPEG | The audio track is extracted, then separated |
The only upload limit is 30 MB per file, which a normal MP3 does not reach until roughly the 30 minute mark, so length is almost always what stops you rather than size. Both outputs are MP3 no matter what went in.
The 30 second limit and what a run costs
Length is measured on the server with ffprobe rather than trusted from the browser, so the real duration of the file decides whether a run is accepted. Nothing in the page can change it.
| Plan | Longest clip in one run | What a 30 second clip costs | Allowance | Results kept for |
|---|---|---|---|---|
| Guest | 30 seconds | Your one free run of the day | 1 run per day per connection | 30 minutes |
| Free account | 30 seconds | 5 credits | 5 credits per day | 2 hours |
| Pro | 600 seconds (10 minutes) | 5 credits | 1500 credits per month | 24 hours |
Billing works in 30 second steps: the cost is 5 credits for every started 30 seconds, so a 30 second clip costs 5, a 31 second clip costs 10, and a 10 minute track costs 100. For a free account that means 5 credits a day buys exactly one 30 second run, with no way to buy half a run or carry unused seconds into tomorrow.
One fair detail: credits are only deducted when a result is actually produced. A file rejected as too long, or a task that fails on the service, costs you time but nothing else, and closing the page mid-run does not spend your allowance.
When your song is longer than the limit
A three minute song does not fit in a 30 second run, and there is no setting that changes that. The workable route is to cut the song down first.
- Trim the section you need. The audio trimmer runs entirely in the browser, so nothing is uploaded while you narrow the track to the 30 seconds you need.
- Separate that section. Upload the trimmed clip and download both tracks as soon as the run finishes.
- Repeat for other sections if you need them. Each pass is a separate run costing another 5 credits per started 30 seconds, and joining the pieces in an editor is your job, because the seams depend on where you cut.
Why even a short clip takes about two minutes
The wait is not your connection. The service carries a fixed start-up cost of roughly 100 seconds before any audio is processed, plus a little more for the audio itself. In our own measurements a 5 second clip completes in about 105 seconds and a 25 second clip in about 125 seconds, so a 30 second clip sits at the slow end of the same two minute band rather than being dramatically worse.
The page polls the job every 3 seconds and shows the elapsed time, so a progress bar that hardly moves is expected, not a sign that something broke. Refreshing does not speed it up: a reload drops the job reference and the result can no longer be collected, so leave the tab open and let it finish.
How the AI separation actually works
The track is sent to Volcengine AI MediaKit, which runs a trained source separation model over the mix. The model estimates two signals, the parts that sound like a human voice and everything else, and writes each out as an MP3. Because it is a prediction rather than a filter, it also works on tracks never released as instrumentals and on mono recordings where centre channel cancellation fails.
The same property explains why the result is not surgically clean. Separation is an estimate, so dense mixes, long reverb tails and distorted vocals are the hardest cases, and small traces of the other side can survive in each file. There is no strength slider to compensate, so if the first attempt sounds rough, try a cleaner source rather than hunting for a setting that does not exist.
What the two tracks are used for
| Task | Track you need | Why |
|---|---|---|
| Karaoke or band practice | Instrumental | A backing track without hunting for an official release that may not exist |
| Remix, mashup or sample | Vocals | A clean acapella drops straight into an editor as a normal MP3 |
| Lyrics or transcription | Vocals | Background music is the most common reason automatic transcription goes wrong |
The tool cannot grant rights, though: copyright in the recording and the composition stays with its owners, so what you may publish or monetise depends on your licence and on the platform rules you post under.
What this tool cannot do
- It uploads your audio. This is a cloud tool, so the file leaves your device and is processed by a third-party AI service.
- 30 seconds is a hard ceiling for guests and free accounts, enforced on the server from the real duration, so a longer clip is refused rather than trimmed for you.
- The free allowance is one run a day at most. A guest gets a single run, and a free account gets 5 credits, which is exactly one 30 second run.
- It is slow. About 100 seconds of fixed overhead plus a little more per second of audio, so a couple of minutes is normal for even a few seconds of sound.
- Two stems only. There is no drums, bass, piano or guitar split. Multitrack stem separation is a different class of tool.
- No adjustable parameters: no separation strength, no choice of parts to keep, and no way to fix a poor result except to change the source and spend another run.
- One file per run, no batch, so a folder of tracks has to be processed one at a time.
- No manual repair: no brush for painting out bleed and no waveform editing, only the two finished files.
- Results expire. 30 minutes, 2 hours or 24 hours depending on your plan, with no later re-download.