Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Reading what it learned

How to tell a real preference from a coefficient that happens to be pointing somewhere.

The TASTE view documents what each tab shows. This page is about reading it well: the interpretation mistakes that are easy to make, and how the interface tries to stop you making them.

Four states, and what each means

The instrument distinguishes four, and never lets two of them look alike:

It saysIt means
not measuredThe feature vector has no coordinate for this. It never will
not fittedNo posterior yet. Answer some duels
too few examplesFewer than five patches in the pool use it. Not enough to fit a coefficient
a value ± an intervalHere is the belief, and here is how much to trust it

A dash is not zero. "The model is indifferent to this" and "the model has never had a chance to form a view" are different statements, and one grey bar cannot say both.

Read the interval, not the bar

The single most useful habit.

In DIRECTIONS, every coefficient is drawn with a credible interval behind it. If the interval crosses the centre line, the model has not established that coordinate: the bar is a guess that happens to point somewhere, and it will likely point elsewhere after ten more duels.

A short bar with a tight interval is worth more than a long bar with a wide one. The former is a small preference the model is sure of; the latter is noise with confidence.

Drag the evidence slider. Early on every interval straddles zero, and the individual bars mean nothing even though they point somewhere. As observations accumulate the intervals narrow and coefficients start clearing zero one at a time. Red whiskers are the ones that have not.

The same logic runs the node bank's θ bars, which is why they draw a dash below five supporting patches. A coefficient fitted from three examples would otherwise look exactly like one fitted from three hundred.

Size on the map is uncertainty

On the MAP, glow is how much it thinks you would like a patch and size is how unsure it is. People read glow and ignore size.

  • Small and bright. Confident it is good. Worth playing.
  • Big and bright. It might be excellent. This is where to explore.
  • Small and dim. Confident it is not for you.
  • Big and dim. It knows nothing. Also worth exploring, for a different reason.

Early in a session everything is big. That is what a cold start looks like, and it is why the first generation you breed is not very targeted.

Also read the variance footer: the two axes typically capture around half the variation in the feature space, so two dots close together are probably similar and two far apart are probably different. It is a projection, not a map of the territory.

Styles are lenses, not genres

A style lens is a direction in feature space that explains some of your answers. It is not a genre and it is not a mood. The generated names (drive & fold + chorus, dynamics + plucked strings) describe coefficients, not music.

Two things follow:

  • A lens claiming almost none of the bank is idle. The model fits up to five and lets the data decide how many get used. Having two live lenses and three idle ones is not a failure; it means your taste, as measured by these coordinates, has two islands.
  • You can rename them, and should. Click a chip's name. Once *"drive & fold
    • chorus"* is "the mean one", every place the style appears becomes readable at a glance. The name is yours and it persists.

The prediction on a bank row

The percentage is roughly "how likely you are to prefer this patch in a duel against an average pool member". It is a posterior mean, so it already accounts for the model's uncertainty by averaging over it, which means a confident 80% and an unsure 80% look identical here.

If you want the uncertainty, that is what the map's size channel and the belief row's interval are for. The row is a ranking aid, not a measurement.

Trust, and what to expect over time

TRUST is the tab that decides whether any of the others deserve belief. A realistic trajectory:

StageWhat TRUST says
First session, < 20 picksNot beating a coin flip. Correct and expected
20–60 picksSkill crosses zero and wobbles. Buckets too small to read
Beyond thatSkill climbs; dots settle near the diagonal

Two failure shapes worth recognising:

  • Dots consistently below the diagonal on the right. It is overconfident: when it says 80% it is right less often than that. Usually a sign it has locked onto a coordinate that was coincidental. More duels, especially ones you expect to surprise it, is the fix.
  • Skill stuck near zero with many observations. Either your preference is not visible in the feature space (see what it cannot learn), or your answers are inconsistent, which happens: some days you are not choosing on one axis.

The number to watch is check-duel skill rather than overall skill. The overall number is measured on questions the model helped choose; the check duels are drawn at random.

Why a low score early is the honest one

Auracle forecasts every duel before you answer it, then reports its own error against a proper scoring rule. A number produced that way can come out badly, and early on it does. That is what makes it worth reading later.

When it is working

You will notice it before the numbers say so:

  • The duels get harder — both candidates are plausible.
  • Generations produce children you want to keep rather than children you want to skip.
  • The belief row's explanation matches your own reason for liking a patch.
  • A style chip's name is one you would have written yourself.

That last one is the real milestone.