You can merge audio files online by adding two or more tracks to a browser tab, setting the order you want, and exporting them as one MP3, WAV, M4A or OGG file. The join runs with FFmpeg compiled to WebAssembly, so the audio is never uploaded, no account is needed and there is no watermark.
Two details are worth knowing before you start. The merger joins tracks end to end, which is the same as concatenation rather than mixing them on top of each other, and every input is resampled to one sample rate and to stereo first, so mixed sources are not stitched byte for byte. A single run accepts up to 20 files and needs at least two. MP3, WAV, M4A, AAC, OGG, FLAC and Opus can all be added, and the first run downloads an engine of about 30 MB.
Joining audio clips is the smallest audio edit there is, and it turns up more often than people expect. A podcast is recorded in three sittings, a meeting arrives as several files, a learner wants one track of repeated phrases. A desktop editor can do it, and it also asks you to install something and learn a timeline you will not touch again for months.
Merge Audio
Add up to 20 tracks, reorder them with the arrows, and join them into one MP3, WAV, M4A or OGG file. The audio never leaves your device.
What merging audio files actually does
An audio merger writes several files into one, in a fixed order. It is not the same job as mixing, which layers tracks so that a voice and a backing track sound at the same instant. Merging, or joining, places one track after another: the end of the first file becomes the beginning of the second.
The merger on this site performs that join inside the browser tab. Each file is read on your device, written to a local working area, joined with the concat filter inside FFmpeg, and returned as one download. Your tracks are never uploaded, and the merge order is simply the order of the list. Keep the two words apart when you search, because a joiner answers "these clips belong one after the other", while a mixer answers "these two sounds should be heard together".
How to merge audio files online
The whole job takes four steps, and none of them require an account or a browser extension.
- Add your audio files. Drop two or more tracks onto the upload area or click to browse. MP3, WAV, M4A, AAC, OGG, FLAC and Opus are accepted, each file is read on your device, and its length appears in the list. Anything that is not audio is refused with a short note instead of being half-processed.
- Set the merge order. The list order is the merge order, read from top to bottom. Use the up and down arrows on each row to move a track, and the trash button to drop one. New files are appended to the end of the list, so the first drop sits at the top.
- Choose the output. Pick MP3, WAV, M4A or OGG. For the lossy formats choose 128, 192, 256 or 320 kbps, with 192 kbps as the default; WAV ignores the bitrate setting because it stores uncompressed samples. Leave the sample rate on Keep original, or set 44100 or 48000 Hz.
- Merge and download. Click Merge audio. The result card shows the segment count, the total length and the size change from source to output, plays the joined file back, and offers the download.
You can hear the result before you keep it. The player on the result card plays the finished file, not a preview of one track, so a wrong order shows up immediately.
Why the merged file is re-encoded, not a byte-for-byte stitch
The intuitive picture of joining two MP3s is that the second file is glued onto the end of the first with the bytes untouched. That is not what happens here, and understanding it explains the quality question.
The concat filter needs every input to share the same audio parameters, so the merger converts each track to one sample rate and to stereo before it joins them. That is what keeps the seams clean: without a common format, a join between a 44.1 kHz mono voice memo and a 48 kHz stereo recording would click, or shift in pitch at the boundary. The price is that the samples are resampled rather than copied, so the output is not bit-for-bit identical to the inputs, even when you export WAV.
| Output | Encoder | Bitrate setting | Best for |
|---|---|---|---|
| MP3 | libmp3lame | 128, 192, 256 or 320 kbps | Playing anywhere, small files |
| WAV | pcm_s16le | Not applicable, PCM is uncompressed | Editing, archiving, further processing |
| M4A | AAC | 128, 192, 256 or 320 kbps | Quality per kilobyte, apps that expect AAC |
| OGG | libvorbis | 128, 192, 256 or 320 kbps | Open, royalty-free playback |
The re-encode is the price of a clean seam. Two inputs that already share a sample rate, a channel layout and a codec are still decoded and encoded again, so even a same-format merge is not byte identical. On speech and podcasts at 192 kbps the audible difference is small enough to ignore. Where it matters is a mastering or archival job, and there the correct answer is WAV for the working copy and one final encode at the end.
When merging audio files is the right job
Podcasts recorded in segments
Recording a conversation in one long take is risky, so many people stop and restart between topics. Every restart leaves a separate file, and publishing means joining them in the right sequence. The merger handles that in one pass, and the arrows let you fix the order before you export.
Meetings captured in several files
Conference phones, mobile apps and voice recorders often split a long meeting at a size or a time boundary, so a ninety minute session arrives as two or three files. Joining them produces one archive copy that is easier to store, transcribe and share than a folder of fragments.
Language practice and shadowing
Learners who stitch a phrase, a pause and the same phrase again end up with a loop they can play on repeat. The same approach works for pronunciation drills and exam listening material that has to run in a fixed sequence.
What this audio merger cannot do
Read these limits before you commit, because in a few cases a desktop editor is genuinely the better answer.
- Two files minimum, twenty files maximum per run. A drop past the twentieth file is truncated and reported as skipped, so larger sets go through in batches.
- The join is a re-encode, so it is not lossless. Every input is resampled to one sample rate and to stereo before the concat, which keeps the seams clean but changes the samples, so a merge of mixed sources is not bit for bit identical even when the output is WAV.
- The first run downloads about 30 MB. That is the FFmpeg engine, which is cached afterwards and then works offline. On a slow connection the first click sits at the loading stage for a while.
- Everything runs in browser memory. All inputs are held at once, because the concat filter needs them together, so a large set can exhaust the tab and fail part way through. Merge in passes when that happens.
- No crossfade between tracks. Tracks are butted together, so a song that ends on a beat jumps straight into the next one. There is no fade, no overlap and no silence insertion, and this is not a DJ mixing tool.
- No loudness matching. If one recording is quiet and the next is loud, the join keeps both levels as they are. There is no auto levelling and no per-track volume, so fix the levels before you merge.
- No trimming inside the merger. The merger only joins. Cut a track first in the audio trimmer, then add the shortened file to the merge list.
Where merging fits in a local audio workflow
Merging is usually one step in a chain rather than the whole task. A common route is to extract the audio, cut it to length, join the pieces, and then convert the result to whatever the destination expects.
The video to MP3 tool pulls the soundtrack out of a video file locally, which is the usual starting point when the audio you want is inside a screen recording or a clip. The audio trimmer cuts a recording down to the part worth keeping, and the audio converter changes the container afterwards if the destination wants a format the merger does not write. The voice recorder captures new audio directly in the browser, so the pieces are already on your machine before they reach the merger.
Audio Trimmer
Cut a recording down to the part you keep, then add the shortened file to the merge list. Also runs entirely in the browser.
Because each of these tools runs in the tab, a file can move through the whole chain without an upload step anywhere.
Why merging locally matters
Most audio on a personal machine is not meant to be public. Voice memos carry family conversations, meeting recordings carry a customer call, and interview tapes carry someone who did not agree to sit on a server you picked. Uploading a set of files to a joiner means trusting a retention policy you have not read.
Running the merge in the tab removes the question. The tracks are read from local storage, decoded and encoded locally, and the only traffic is the page itself plus the one-time engine download. You can check that rather than take it on faith: open the developer tools, switch to the Network panel and run a merge.