Research record — Published from the project repository. Each record states its own date, scope and evidence status.

Research records: Study protocol

Musical utility: study protocol and evidence requirements

Discussion protocol · 13 September 2026. Companion to the research agenda, Antibes study, and Bruckner study. These experiments have not been run. No sample size, success threshold, interface, or sensor choice below is an established result.

1. Separate the questions

The research has four distinct aims:

  1. Realization: help a performer produce a musical result they already have in mind.
  2. Control: help a performer make and repeat a chosen change.
  3. Exploration: help a performer discover and develop worthwhile material through playing.
  4. Learning: help a performer retain useful relationships and act on them later.

A pleasing generated result can support exploration without demonstrating realization. Faster completion can support control without establishing expressiveness. A beautiful display can be enjoyable without improving learning. Report outcomes separately.

For private musical imagination, there is no directly observable ground-truth waveform available to the experimenter. Use the performer’s prior description, sung or played sketch where feasible, successive judgments, and retrospective account as complementary evidence. None alone is an exact measurement of their interior experience. Keep tasks with an external target separate from tasks originating in imagination.

2. Establish source and version provenance

For every recording, record performer or conductor, ensemble, work, recording date and venue where known, release identifier, edition, track titles, and access source. For scores, record composer, work, movement, version, editor, publisher, and page/bar-number convention. A performance can depart from its stated score, so record cuts or other discrepancies independently.

For Antibes, use the named Verve track block or an identified edition containing the same concert. Keep track-local timestamps authoritative for navigation; cumulative timings from metadata are provisional. Do not reuse timestamps from Porter’s incomplete tape or assume the surviving film covers the full suite.

For Bruckner, start from the dossier’s declared 1890/Nowak Eighth and 1878/Nowak Fifth references, subject to the founder’s preferred recordings. If a chosen performance uses Haas, an earlier version, or an unidentified text, label that explicitly and revise the correspondence plan. The first phase should not conflate interpretation differences between conductors with differences between score versions.

No full copyrighted audio, scans, or third-party transcriptions need be committed to the repository. Store citations and annotations; use lawful local access for listening. Use original miniatures or appropriately available material for redistributable prototype demonstrations.

3. Preserve evidence before extracting features

Retain the original source identifier and available performance data. Derived records should retain event identity, onset, duration, absolute pitch or pitch trajectory, voice/instrument, articulation, dynamics, and relevant uncertainty. Simultaneous and overlapping events must remain representable. Retain rests and phrase boundaries without inventing cyclic closure.

Distinguish three levels of timing: recording time, notated score time, and a derived beat grid. Rubato, pauses, and asynchronous voices make them different. Annotators should be able to record a boundary interval rather than a falsely precise instant.

For each analytical view, declare its reduction: octave removal, pitch quantization, deduplication, voice selection, onset quantization, window choice, or cyclic closure. Keep the source available. The existing RMCP converter is useful for its stated analysis workflow; it is not suitable as the sole performance archive for these studies. In particular, its rational pitch lookup is not a preservation of recorded intonation.

Record an explicit distinction between observation, source attribution, interpretation, and experimental instruction. “Piano onset at approximately t” and “pianist responds to saxophone” have different evidential status. A source-backed claim about a studio recording does not become an observation of a live version.

4. Listening before model fitting

The initial researcher/performer pass should be uninterrupted and should allow free descriptions, humming, gesture, or sketches. Ask where an idea returned, where attention moved, and where a continuation or ending became plausible. Do not initially require the terms symmetry, tension, sacredness, fractal, or climax.

Afterward, select a small number of passages that test specific questions. Include relatively sparse passages, transitions, pauses, and endings as well as peaks. Use the case dossiers to choose productive targets, while allowing the listening to contradict the preliminary plan.

On a second pass, annotate relevant voices independently. Have a second listener annotate a subset before reconciliation. Preserve disagreement and uncertainty. Automatic features can locate candidates; pitch tracking errors, masking, tape artifacts, and source separation errors must not become unnoticed musical conclusions.

A minimal annotation record should include source/version, movement, local time interval, any score locator, observed event or feature, analytical interpretation, uncertainty, annotator, and evidence category. Do not present this suggested record as a newly implemented repository schema.

5. Initial making experiment

The first proposed instrument is the phrase-transformation concept in the agenda. Begin with a phrase small enough to remember and repeat. Choose at most two or three operations after the performer identifies the most relevant intention and input modality. A fixed symbolic transposition may be an initial operation; expressive audio transformation is a different technical scope.

Use both an externally specified target task and an original-idea task. For the latter, ask the performer to capture or describe the intended phrase before manipulating it where that is feasible, then permit changes of intention and record them. Intention need not remain frozen throughout creative work.

Compare the experimental interface with the performer’s familiar tool, with equal practice time and counterbalanced order. Include a simple interface exposing the same operations where possible; otherwise improved results could reflect a larger feature set rather than the geometry. Keep synthesis and audio quality comparable when they are not the manipulated variables.

Allow a short pilot to reveal unsuitable tasks and controls. Choose primary outcomes and decision criteria after the pilot and before a confirmatory study. An initial study with the founder can guide design but cannot establish general usefulness for musicians. Later participants should be recruited for the specific task: experience with improvisation, orchestration, or a controller can materially change the comparison.

6. Candidate experiments and interpretable outcomes

P1 — Repeatable transformation

Ask the performer to realize a named transformation, return to the original, and repeat the transformation. Record task time, corrections, unintended changes, and whether the preserved properties really remained fixed. Then ask whether the control felt musically useful. A mathematically correct operation with poor usability is a meaningful negative result.

P2 — Recognizable variation

Let the performer change timing, register, articulation, or pitch order independently using an original motif. Collect identity judgments from the performer and separate listeners. Compare full-event, content-only, and order-aware views only where they expose comparable controls. Do not label every valid permutation a faithful variation.

P3 — Return after intervening material

Present or perform a phrase, introduce other material, then return with selected changes. Compare local similarity measures with recognition and accounts of the return’s role. This is the shared question motivated by Antibes and the Eighth Adagio; it does not presume they use the same compositional mechanisms.

P4 — Independent voices and contingent response

For the Fifth-inspired task, ask the player to keep two ideas perceptible while controlling their entries and prominence. For the quartet-inspired task, compare accompaniment responding to the player with a nonresponsive replay matched on basic activity. These are different experiments. Neither a listener’s preference nor a temporal correlation in a historic recording proves causal interaction.

P5 — Visual transfer and divided attention

Compare performance with and without the visualization after equal training. Include a delayed task without the display. Measure whether the player retains control or recognition, alongside visual attention demands and reported distraction. Assess audiovisual performance as a separate artistic use where the display is part of the work, not an aid expected to disappear.

P6 — Personal mapping and stability

Compare a performer-taught mapping with a fixed mapping for the same operations. Test nearby gestures, repeatability, recovery, and retained control after a delay. Keep adaptation frozen within a measured performance block. Ask whether changes reflect increased skill, model updates, or both. Preserve the ability to restore an earlier mapping.

7. Measurement without reducing music to one score

Collect technical timing, task behavior, and first-person accounts separately. Measure end-to-end gesture-to-sound latency and its variation on the actual hardware; theoretical compute speed is insufficient. Report median and upper-tail behavior, dropped or stuck events, and the timing of interruptions. Set acceptable targets in relation to the musical task rather than announcing an unsupported universal cutoff.

Behavioral measures can include repeatability, number of corrective actions, time to a satisfactory result, and recovery after an unintended change. Recognition and perceived control can be gathered through appropriately designed ratings and interviews. Longitudinal evidence can include voluntary return to the instrument, repertoire developed for it, and examples of ideas made possible by sustained practice.

Keep liking, novelty, agency, fidelity to intention, and quality of finished work separate. Qualitative evaluation of expressive interfaces has a precedent in the NIME work of Stowell, Plumbley, and Bryan-Kinns. Fiebrink, Cook, and Trueman also show why user judgments of interactive models need not track conventional prediction accuracy.

Small exploratory samples should yield descriptive results and concrete examples, not broad population claims. For a later comparative study, base sample planning on its specific primary outcome and pilot variability. Preserve unsuccessful trials and explanations that compete with the preferred hypothesis.

8. Representation and proof checks

Before evaluating a musical operation, check that its advertised properties hold on its declared domain. Candidate checks include source-data retention, event-level round trips, transposition composition/inverse laws, timing preservation under a pitch-only operation, voice-identity retention, and recovery through recorded history. These are proposals for future checks, not additions to F01–F10.

When an operation leaves the discrete musical seed, document the realization rule. If a continuous path is quantized, measure the resulting changes and verify only the properties that survive. Distinguish a proof about a model from correspondence to running software and from an empirical claim about a performer.

9. Exit criteria for choosing a first implementation

The discussion should select one user, one recurring musical task, one input modality, and a small operation family. A useful proposal must identify a baseline, the intended improvement, the information it retains and discards, and an outcome that would count against it.

A prototype earns further work when repeated use reveals a specific expressive or learning benefit and a manageable failure pattern. Attractive geometry, increasing mathematical complexity, or a high detector score alone does not meet that criterion. A negative result should be allowed to narrow the project or move the interface toward a different modality.

10. Work completed and work still to perform

The repository now contains the conceptual agenda, source-based case studies, verified edition metadata, and this proposed protocol. The present task did not produce direct listening annotations, new transcriptions, a complete score audit, or an instrument prototype. Those are the next empirical steps after recordings, editions, and the performer’s priorities are selected.