What to Send Your Mixing Engineer
Most delays in a mixing commission are not caused by difficult music. They are caused by files — the wrong format, effects that cannot be undone, or tracks that do not line up. Every one of these is avoidable in about ten minutes before you send anything.
This guide covers exactly what to send, in what format, and how to read the Japanese terminology you will run into if you commission an engineer in Japan. If you are working on a Japanese song cover and have not sorted out your instrumental yet, start with where to find official off-vocals first.
- The three files you always send
- Japanese mixing terms, decoded
- Export settings that actually matter
- The single most common mistake
- Why every file must start at the same point
- Harmonies, doubles and ad-libs
- Reference tracks — the cheapest thing you can send
- Pre-send checklist
The three files you always send
| # | What | Details |
|---|---|---|
| 1 | Your vocal, unprocessed | WAV, 24-bit, no effects of any kind. Harmonies and ad-libs as separate files. |
| 2 | The instrumental | The exact file you recorded against, at the best quality you have. If you only have an MP3, say so. |
| 3 | A reference | A link to a finished song that sounds like your goal. One is enough. |
Anything beyond these three is a bonus, not a requirement. If you have a lyric sheet with the sections marked, or notes about which take you prefer at a particular line, send them — but do not delay the project gathering extras.
Japanese mixing terms, decoded
If you commission an engineer in Japan, some of these will appear in the brief. They are not interchangeable, and confusing two of them is the usual cause of a wasted round trip.
| Term | Japanese | What it actually refers to |
|---|---|---|
| inst | インスト | The backing track with no lead vocal. What you sang along to. |
| off-vocal | オフボーカル | The same thing. Preferred term in Vocaloid circles. |
| dry | ドライ音源 | Your voice with no effects at all. This is what you send. |
| 2mix | 2ミックス | Ambiguous. Can mean a stereo instrumental, or a rough mix of vocal plus instrumental. Always confirm which. |
| para | パラ | Separated tracks. In cover work it usually means your vocal parts split into individual files. |
| a cappella | アカペラ | Vocal only, no backing. Sometimes requested as a deliverable, not as a source. |
Export settings that actually matter
| Setting | Use | Why |
|---|---|---|
| Format | WAV | MP3 permanently discards high-frequency detail that de-essing and pitch correction depend on |
| Bit depth | 24-bit | More headroom for correction. 16-bit works but gives away margin for nothing |
| Sample rate | Match your project (usually 48 kHz) | Converting adds a processing step with no benefit |
| Effects | None | Cannot be removed later. See below |
| Normalisation | Off | Changes level without improving anything, and can disguise clipping |
| Peak level | Roughly −12 to −6 dB | Comfortable. Anything touching 0 dB may already be clipped |
| Start point | Beginning of the project | Keeps every file aligned. See below |
| File names | Descriptive | MainVocal.wav, Harmony_High.wav — not Audio 1.wav |
The single most common mistake
Sending vocals with effects already printed into them. It is the most frequent avoidable problem we see, and it is almost always well-intentioned — the singer added reverb because it sounded better while recording, then exported without removing it.
Here is why it matters. Mixing largely consists of shaping a raw signal: controlling dynamics, removing resonances, correcting pitch, then placing the voice in a space. Effects printed into the file sit in front of every one of those steps, and cannot be undone.
| If this is printed in | What it prevents |
|---|---|
| Reverb / delay | The tail is baked into the sound. Any new reverb stacks on top, and the vocal cannot be moved forward in the mix |
| Compression | Dynamics are already flattened. Quiet detail cannot be recovered |
| Autotune / pitch correction | Artefacts are permanent, and further correction compounds them |
| EQ | Removed frequencies are gone. Boosting them back raises noise instead |
| Noise reduction | Aggressive settings thin the voice in ways that cannot be reversed |
If you genuinely want a specific effected character, send the clean take and a rough version with your effects as a reference. Then the sound is communicated without being locked in.
Why every file must start at the same point
Export every track from the very beginning of the project — timeline zero, or bar one — not from where the singing starts.
The reason is simple: files exported this way all have the same length, so they drop into any session and line up automatically. Files trimmed to their first note each have a different offset, and the engineer has to align them by ear. On a track with several harmony layers this is slow, and every manual alignment is a chance for a timing error.
Your DAW will call this something like “export from project start” or “whole session length”. Leading silence is not a problem — it is the point.
Harmonies, doubles and ad-libs
Keep them separate. One file per part, all starting at the same point.
The reason is control. Harmony sits behind the lead; doubles widen it; ad-libs need to appear and disappear. If they are bounced together into one file, none of those decisions can be made independently, and the engineer is left balancing choices you already made permanently.
A workable naming scheme:
01_MainVocal.wav02_Harmony_High.wav03_Harmony_Low.wav04_Double.wav05_Adlib.wav
Numbering matters more than it looks — it keeps the order intact when files are unzipped on a different operating system.
Reference tracks: the cheapest thing you can send
A reference is a finished, released song that sounds roughly like what you want. It is the highest-value item on the list, and it costs you a single link.
Descriptive words do not survive translation between people. Clean, warm, powerful, natural — every engineer has heard all four used to describe opposite things. A reference removes the ambiguity completely.
How to choose one well:
- Pick a similar genre and vocal type. Referencing a heavily produced pop record for a quiet acoustic cover does not transfer.
- Say what specifically you like about it. “How present the vocal is in the chorus” is far more useful than “this vibe”.
- One is enough. Three references pointing in different directions is worse than none.
Pre-send checklist
| Check | If not | |
|---|---|---|
| 1 | All vocals exported as WAV, 24-bit | Re-export. Do not convert an MP3 to WAV — the loss already happened |
| 2 | No effects printed in | Bypass everything on the track and export again |
| 3 | Every file starts at project start | Re-export using whole-session length |
| 4 | Harmonies and ad-libs are separate files | Export each part individually |
| 5 | Instrumental is the exact file you sang to | Find it. A different master will not line up |
| 6 | Nothing peaks at 0 dB | Check for clipping and re-record if the take is distorted |
| 7 | Files are clearly named and numbered | Rename before zipping |
| 8 | One reference link included | Add one. It takes a minute and changes the outcome |
Unsure whether your export is clean? Send one file and we will check it before you do the rest.
We mix utaite, youtaite and VTuber covers from Tokyo — all communication in English by email, X or Discord.
One last thing
None of this is about being technical for its own sake. Every item above exists to protect a decision that is still reversible. A clean 24-bit file starting at bar one keeps every option open; an effected MP3 trimmed to the first note has already closed most of them.
You only have to set this up once. After the first project the export settings become a preset, and you never think about it again.
Next: if you are still assembling the pieces of your cover, see the complete guide to making an utaite-style cover, or how to find official off-vocals for Japanese songs.