What AI mastering is good at, and what it flattens
An honest account of where the automatic services work, where they do not, and how to tell which situation you are in before you pay for either option.
Automatic mastering services are now good enough that pretending otherwise is not useful to anyone. They are also not equivalent to a person, in ways that are specific and predictable rather than mystical. Knowing which job you have is worth more than an opinion about the technology.
What they do well
The underlying task — measure a file, compare it to a model of finished records, apply corrective EQ and dynamics, land on a sensible level with a safe peak ceiling — is genuinely well suited to automation. It is measurement and correction, and machines are reliable at both.
So on material that resembles the model, the results are good. A conventionally arranged song with a clear balance, no unusual dynamics and nothing eccentric in the low end will come back louder, tonally tidier and technically compliant. For a demo, a reference, a playlist submission or a release with no budget, that is a real service at a price that makes a human pass impossible to justify.
They are also fast and consistent, which matters more than it sounds. Twelve tracks mastered by the same process in an afternoon will at least be consistent with each other, and consistency across a set is a problem that catches a lot of self-released records.
For conventional material on no budget, an automatic master is a real answer, not a compromise you should feel bad about.
Where they fail
The failures are not random. They cluster around material that departs from the model, and they take one of three shapes.
- Unusual dynamics get flattened. A record built on the contrast between a whispered verse and an enormous chorus is, to a measuring system, a record with inconsistent loudness. The correction reduces the contrast, which is the thing the song is made of.
- Deliberate tonal choices get corrected. A mix that is intentionally dark, intentionally thin, intentionally bright as a stylistic position reads as a deviation from the average and gets pulled toward the average.
- Unusual low end gets misjudged. Sub-heavy material, or arrangements where the bass and kick occupy an intentionally strange relationship, tend to come back either weakened or inflated, because the model has an opinion about how much energy belongs down there.
The common thread is that these systems are trained toward a centre. That is the correct design — the centre is where most material should go — but it means the further your record sits from the middle, the more the process treats its character as an error.
The reference problem
Most services accept a reference track and claim to match it. The matching is real but it is shallower than the word suggests.
What gets matched is measurable: spectral balance, loudness, dynamic range, stereo width. What does not get matched is everything that made the reference sound the way it does — the arrangement, the performances, the mix decisions, the specific way one engineer handled that particular snare.
Feeding in a reference whose mix is fundamentally different from yours does not move you toward it. It applies a tonal difference curve between two unrelated records, which frequently makes things worse. References work best when the two mixes are already close — which is exactly when you need the help least.
How to tell which you need
Four questions, answerable in a minute.
Does the record depend on dynamic contrast? If the emotional mechanism of the song is loud against quiet, a measuring system is the wrong tool, because reducing that variation is the specific thing it does.
Is anything about the mix a deliberate departure? Intentionally dull, intentionally harsh, intentionally unbalanced. If yes, the correction will fight you, and you will spend longer arguing with presets than a person would have taken.
Is this a set that has to hold together? Automatic services handle per-track consistency reasonably and sequencing badly, because sequencing is a set of musical judgements about a listening experience, not a set of measurements.
And: is anyone paying for this? A release with a budget, a campaign, a deadline and someone's reputation attached is not the place to save two hundred pounds. A track going out on a Thursday to four hundred listeners is.
Ask what the record depends on. If it depends on something a measurement cannot see, a measurement cannot protect it.
Checking the output
Whichever route you take, the same verification applies, and it is the part people skip on automatic masters because the service implies it has been handled.
Level-match the master against your mix and compare. You are listening for what was lost, not for what was gained — the gain is obvious and the loss is not. Check that the quiet sections are still quiet relative to the loud ones. Check the low end has not been inflated. Check true peak is below minus one. Check sibilance survived.
If the master is louder and the record is smaller, the process did its job and the job was wrong for this material. That is not a failure of the technology, it is a mismatch, and the answer is either a different setting or a different route.
The tools are good. They are tools. The judgement about whether this record wants what they do is still yours, and it is the part that has never been automatable because it is a question about intent rather than about signal.
Keep reading
What to charge when you have no idea what to charge
A method that does not require knowing what anybody else charges, which you mostly cannot find out anyway.
Read →Two revisions, in writing, every time
The single clause that prevents most of the disputes in this trade.
Read →Nobody agrees what a stem is
Ask five engineers for stems and you will get five different things. Here is how to specify it so the wrong one never arrives.
Read →All notes
Every note, newest first.
Back to Notes →