hn.today

Learning Jazz Pianist Style with Cross-Attention Conditioning

almostimplemented.github.io23 points2 comments
Screenshot of Learning Jazz Pianist Style with Cross-Attention Conditioning

This describes a system that learns to generate solo jazz-piano continuations in the stylistic manner of twelve labeled pianists by fine-tuning Aria, a 16-layer transformer pretrained on piano MIDI. The authors insert gated cross-attention blocks into the last eight layers so the model attends at every step to a small learned embedding (four vectors) for the chosen pianist; a learned gate (initial value 0.1) controls how much the conditioning alters the residual stream. Training uses automatic-transcription MIDI from the PiJAMA dataset and prompts of 256 tokens to produce 4,096-token continuations; MIDI rendering in the demos is an approximate sampled piano.

Style is evaluated not by perplexity but by who a strong sliding-window pianist classifier hears. Perplexity barely changes (11.41 → 6.96 → 6.82) while classifier agreement with intended pianist rises from 25% (pretrained Aria) to 37% (fine-tuned, no conditioning) to 70% (fine-tuned with conditioning); chance is 8%. A classifier trained only on generated music transfers to real recordings (87% of chunks, 95% of songs) and a separate analysis locates “where style lives” within performances. Results vary by pianist (e.g., 96% for Hank Jones and Dick Hyman vs 29% for Cedar Walton). Outputs are selected MIDI samples, not copies of training data; limitations include MIDI/transcription fidelity and reliance on an automated classifier rather than a listening study.

Read on almostimplemented.github.io2 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.