ConceptioArchivearXiv CS
arXiv CSopen access

ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation Jhen-Ke Lin National Yang Ming Chiao Tung University [email protected]

MUSICAL FIT

PLAY-RELEVANT CHARACTERISTICS

A

B

A

C

TIMING

MUSIC RESPONSE

NOTE SEQUENCE

REPETITION & FORM

HUMAN-CHART DISTANCE

DIFFICULTY & LIMITS

On the musical grid?

Follows musical activity?

Local note patterns?

Structure across sections?

Near same-level charts?

Typical and feasible demands?

arXiv:2607.12857v1 [cs.SD] 14 Jul 2026

Figure 1: The six dimensions of ChartGenEval. ChartGenEval organizes chart evaluation into six distinct questions, grouped by musical fit and play-relevant characteristics. Each dimension retains its own interpretation.

Abstract

Agreement with one official chart is therefore useful but incomplete. A placement F1 score measures reconstruction: a valid alternative loses credit in the same way as an error. Other automatic proxies answer narrower questions. Language-model (LM) likelihood measures model fit, while diversity and self-similarity describe generated material. None alone tells a developer which chart property changed or how to improve the generator. Fast model iteration needs automatic feedback that keeps these questions separate. Our key observation is that freedom of note choice does not remove external constraints. A chart may choose what notes to place, but it cannot choose the song’s musical time. We therefore use the matched official chart only as an authored map of bars, meter, and tempo. Its note placements never become targets. This separation preserves alternative designs while exposing chart-wide timing shifts that interval- and type-based statistics cannot see. ChartGenEval organizes evaluation into the six questions in Figure 1. Its corruption-tested core covers timing, local note sequence, local repetition, and difficulty or playability limits. Each output has a role: a directional signal can guide optimization, a feasibility check can impose a constraint, and a descriptive value can monitor model behavior. The profile does not collapse these roles into one score. We test the core outputs by injecting known failures at increasing strengths. A useful output should react to its target failure and remain unchanged under predeclared invariance controls. Seven output axes meet these criteria in nine nonredundant tests on 80 held-out song groups. Complementary development stress tests yield two broader lessons. First, timing must be anchored to the song: our phase estimate recovers 15–60 ms chart-wide shifts that chart-only statistics miss. Second, a plausible proxy can reward the wrong behavior: common-pattern rewriting improves perplexity, while loop collapse increases self-similarity.

A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Reference-note agreement therefore measures reconstruction, not the full design problem. We introduce ChartGenEval, a six-question evaluation framework with an automatic, corruption-tested core. It leaves note choice open while anchoring timing to the song: the matched official chart supplies only its authored timing map, never target notes. We test each core output with dose-controlled failures rather than assume that a familiar statistic measures chart quality. Across 80 held-out song groups, seven output axes satisfy prespecified sensitivity and invariance criteria in nine nonredundant tests. Complementary stress tests on the 40-song development panel expose two broader lessons. A chart-wide phase estimate recovers injected shifts of 15, 30, and 60 ms while chart-only outputs remain essentially unchanged. Common-pattern rewriting lowers mean language-model perplexity by 37%, and loop collapse raises mean self-similarity by 62%. ChartGenEval therefore reports separate, role-specific signals instead of one proxy or total score. This profile provides automatic feedback for comparing and iterating generators; selected outputs are candidate optimization targets or constraints after task-specific stress testing.

1

Introduction

Chart evaluation is difficult because charting is a design task: multiple timed, typed note sequences can fit the same song and difficulty. Automatic generation has progressed from onset detection to conditional sequence models [1–4], but evaluation must still distinguish valid alternatives from errors.

1

Our contributions are:

LM loss is retained only as a description of model fit and as a stress-test baseline. In particular, an unseen transition is unfamiliar to the training corpus, not necessarily invalid.

1. Automatic, dimension-specific feedback. ChartGenEval separates evaluation questions and output roles, allowing generators to be compared and revised without treating one reference note sequence as the only valid answer.

3.2

Several chart properties have no universal “more is better” direction. A chart can be too sparse or too dense, and too little or too much repetition can both be undesirable. We therefore measure departure from the middle 80% of same-course human charts. For a feature value 𝑥, let 𝑙 and 𝑢 be the 10th and 90th training percentiles and ℎ = max((𝑢 − 𝑙)/2, 10−9 ). The band score is

2. Corruption-tested signals and reusable diagnostic lessons. Nine nonredundant held-out tests support seven core output axes. Complementary development tests identify a translation blind spot in chart-only timing statistics and wrong-way incentives in perplexity and self-similarity.

2

  (𝑙 − 𝑥)/ℎ,    𝑑 (𝑥) = 0,    (𝑥 − 𝑢)/ℎ,  1 . 𝑠(𝑥) = 1 + 𝑑 (𝑥) 2

Related Work

Reference-chart agreement measures reconstruction against one authored sequence. DDC and GenéLive! report F-scores for predicted placements, while TCP adds onset-plus-type F1 and local pattern recall [1, 3, 4]. TaikoNation compares binary timing frames and pattern spaces [2]. These measures remain valuable for reconstruction, but they treat one authored sequence as the answer. Unpaired statistics and human studies answer complementary questions. DDC uses LM perplexity and token accuracy to measure held-out model fit [1]. Symbolicmusic and text-generation work adds diversity, distribution, self-similarity, and repetition statistics [5–9]. Human and expert studies instead evaluate system-level experience [3, 10, 11]. ChartGenEval complements these approaches by testing whether automatic outputs respond to specified chart failures. Metric stress tests show that a plausible statistic can miss a musical failure. SCHmUBERT constructs nonmusical piano rolls that match framewise self-similarity statistics [12]. Fréchet Music Distance is tested with pitch and velocity perturbations, but remains a set-level embedding distance [13]. STRUM reports reference-onset F1 after a per-song global-offset search [14]; ChartGenEval reports the offset itself without using reference notes as targets. Its controlled corruptions test both target response and invariance, turning metric choice into an empirical question.

3

Methods

3.1

Evaluation questions and core outputs

Same-difficulty calibration

𝑥 < 𝑙, 𝑙 ≤ 𝑥 ≤ 𝑢, 𝑥 > 𝑢,

(1)

Thus 𝑠(𝑥) = 1 inside the band and decreases symmetrically outside it. This is a difficulty-conditioned typicality reading, not a verdict on creativity or overall quality. One-sided limits, such as the density-spike check, remain separate because a severe local overload should not be canceled by another high score. The three less familiar outputs use simple transformations. If 𝑟 is the rate of interval–type trigrams unseen in same-course training charts, transition familiarity is exp(−𝑟/0.08). For 𝑀 4-gram windows with 𝑈 unique windows, the raw repetition rate is (𝑀 − 𝑈)/𝑀 before band calibration. The density-spike score is exp(−𝑒/1.5), where 𝑒 is the amount by which a chart’s 95th-percentile local density jump exceeds the same-course limit.

3.3

Data and timing anchors

We separated design from confirmation. The seven core outputs and their corruptions were designed on 40 development song groups with 170 charts, then tested on 80 held-out groups containing 333 charts. The calibration set contains 3,880 course charts from 813 separate song groups. Songs, rather than courses or corruption replicates, are the inferential units. Appendix A gives the corpus construction, completeness rules, and full analysis protocol. Timing requires a reference outside the generated event sequence. Our primary grid comes from the matched official file’s authored bar timestamps, meter, and tempo. We never use its note times or note types as targets, and we never accept the grid declared by the generated chart. The first choice preserves freedom of note design; the second prevents a generator from redefining the clock against which it is evaluated. An audio-estimated alternative is reported in Appendix A.

ChartGenEval treats the six dimensions in Figure 1 as separate questions. The main analysis focuses on seven outputs for which we have held-out controlled-corruption evidence (Table 1). Timing uses three readings because dense noise, rare outliers, and a shared offset are different failures. The other outputs cover local transitions, local repetition, requested note rate, and density spikes. All outputs are automatic, but automation does not make them interchangeable. Grid-phase error and timing tails have a clear direction; density spikes provide a courserelative density constraint; transition familiarity and typicality are safer as soft targets or monitoring signals. Raw

3.4

Measuring timing against the song

A common shift 𝛿 leaves every inter-note interval unchanged: (𝑡 𝑖+1 + 𝛿) − (𝑡𝑖 + 𝛿) = 𝑡𝑖+1 − 𝑡 𝑖 .

2

(2)

Evaluation question

Automatic reading

Present evidence

Timing Note sequence

C1, C1s, and C2 held-out tests C3 held-out test

Repetition and form

Timing clean rate, timing-error p99, and whole-chart grid-phase offset Transition familiarity: a one-sided score derived from the rate of unseen interval–type trigrams Same-course typicality of repeated versus unique 4-grams

Response to music

Bar-scale energy response and onset support

Human-chart distance

Distance from the same-course training distribution

Difficulty and limits

Same-course note-rate typicality and a local density-spike limit

C4, C5, and C8 test local repetition; longform readings remain exploratory Exploratory development analysis; Appendix C Exploratory development analysis; Appendix C C6 and C7 held-out tests

Table 1: Evidence differs across the six questions: seven core outputs have held-out tests, while the response-to-music, human-distance, and long-form readings remain exploratory. The note types are also unchanged. Chart-only measurements built from these intervals and types therefore cannot identify the shift. This translation symmetry explains the blind spot before any experiment is run. Local timing still needs both a common-error summary and a tail summary. We treat a note within 6 ms of the authored grid as locally aligned, at the scale of published temporal discrimination for short intervals [15]. Timing clean rate summarizes common errors; the 99th percentile of absolute error exposes rare large errors that a mean can hide. Unmatched generated notes remain in the denominator, so timing clean rate must be read with note rate. Full matching rules and the 6/12/18 ms diagnostic ranges appear in Appendix A. Nearest-grid matching can miss a constant shift (Figure 2). On a dense grid, each shifted note may simply be paired with another nearby grid point. We therefore estimate one offset for the whole chart. Let 𝐺 𝑘 be the fixed bar, beat, and eighth-note grids, and let 𝑑 (𝑥, 𝐺 𝑘 ) = min𝑔∈𝐺𝑘 |𝑥 − 𝑔|. For note times 𝑡𝑖 , we score a candidate shift by   ∑︁ 𝑑 (𝑡𝑖 − 𝛿, 𝐺 𝑘 ) 1 ∑︁ 𝑆(𝛿) = 𝑤𝑘 1− , 𝑛 𝑖 𝜏 + 𝑘 (3) 𝛿ˆ = arg max 𝑆(𝛿),

strongest-dose contrast against the matched control, and every relevant invariance check to pass. We use songcluster bootstrap intervals adjusted across ten prespecified rows. Full decision rules and sample accounting appear in Appendix A.

Results

4.1

Core outputs respond to target failures

Every core output has the expected negative dose-rank association under its target failures (Table 2). The result covers nine nonredundant corruption–measurement pairs but seven output axes because the same 4-gram axis detects three distinct edits. All applicable invariance controls pass. The important result is not that every corruption needs a new metric. A small set of interpretable outputs can cover different failure mechanisms when each output is tested against a concrete claim. Sparse errors show why a tail statistic is needed when most notes remain correct. Complete sample counts and exclusions appear in Appendix A.

4.2

External time detects chart–audio shifts

Translation symmetry predicts that a chart can look internally plausible yet be late relative to the song. The 40-song development experiment confirms this mechanism. After shifting every note by 15, 30, or 60 ms, no primary chartonly output changes by more than 0.03 human standard deviations. The phase estimator’s median matches each injected shift. With either the authored or audio-estimated grid, at least 169 of 170 charts are within 3 ms of the injected value. Local matching and global phase answer different questions. A dense grid can re-pair a shifted note to another subdivision and still report a small local error. The wholechart estimate preserves the shared offset instead. The same principle explains sparse failures: the mean remains dominated by correct notes, whereas timing-error p99 exposes the tail. Timing evaluation therefore needs both a song-side anchor and summaries at more than one scale.

𝛿 ∈ [ −250,250] ms

where [𝑧] + = max(𝑧, 0), 𝜏 = 12 ms, and the bar, beat, and eighth-note weights are 0.2, 0.3, and 0.5. We search in 0.5 ms steps. Combining three grid levels reduces periodic ties. The resulting 𝛿ˆ is a chart-level diagnostic: its magnitude says how far the chart must move to best align with the fixed musical-time framework. BPM alone provides a period but not a phase, so it cannot replace this anchor.

3.5

4

Controlled corruptions and acceptance rules

Starting from human charts, each corruption strengthens one operational failure. Other outputs may also move, so the tests establish target sensitivity rather than exclusive specificity. The target output and expected direction were fixed before the held-out run; all transformations appear in Appendix B. After orienting outputs so larger is better, support requires a negative dose-rank association, a negative 3

LOCAL MATCHING

WHOLE-CHART OFFSET

grid agreement

estimate: +60 ms

fixed grid

notes +60 ms

0

100

200

300

400

500

−100

time (ms)

−50

0

50

100

candidate chart offset (ms)

Figure 2: Why timing needs a whole-chart estimate. Left: nearest-grid matching can pair shifted notes with different subdivision points and report small local errors. Right: a schematic agreement curve over candidate whole-chart offsets; the curve peaks at the injected +60 ms shift. Local plausibility does not imply correct song alignment. Targeted failure

Core output

Dose trend [99.5% CI]

Strongest-dose change [99.5% CI]

Dense timing jitter Sparse 60 ms timing errors Whole-chart shift Note-type shuffle Loop collapse Common-pattern rewrite Note-rate scaling Local burst insertion Bar-order shuffle

timing clean rate timing-error p99 grid-phase offset transition familiarity 4-gram typicality 4-gram typicality note-rate typicality density-spike limit 4-gram typicality

−.744 [−.815, −.667] −.234 [−.344, −.134] −.916 [−.981, −.826] −.800 [−.856, −.737] −.445 [−.573, −.312] −.358 [−.459, −.258] −.693 [−.784, −.579] −.940 [−.950, −.930] −.589 [−.728, −.435]

−.502 [−.558, −.446] −.306 [−.508, −.144] ms −55.58 [−61.43, −49.83] ms −.231 [−.262, −.199] −.165 [−.226, −.108] −.131 [−.189, −.078] −.582 [−.672, −.483] −.619 [−.644, −.593] −.224 [−.293, −.156]

Table 2: Held-out results on 80 song groups. Outputs are oriented so larger is better; a useful response is therefore negative as a failure becomes stronger. Every row meets both criteria, and all applicable controls pass. Intervals are family-wise 99.5% song-cluster bootstrap intervals. The two originally named 4-gram variety and repetition scores are algebraic mirrors, so they appear once as one effective axis.

4.3

Common proxies can reward the wrong behavior

C5 common-pattern rewrite LM loss (lower looks better)

Two familiar proxies move in the wrong direction under targeted failures (Figure 3). LM perplexity rewards commonpattern rewriting: on the development panel, mean perplexity falls by 37%, from 9.44 to 5.98. A lower-is-better reading therefore prefers the corrupted chart, opposite to the intended property [16]. The reversal remains with a separately fitted scoring LM, while 4-gram typicality moves away from the human range. Secondary checks appear in Appendix C. If self-similarity is maximized as a quality objective, it rewards loop collapse. Repeating the dominant 4-gram raises mean self-similarity by 62% even though the notetype stream has collapsed into a four-event cycle [9]. 4gram typicality moves in the opposite direction. The practical rule is to stress-test a candidate reward against the failure it is meant to prevent. Automatic profiles shorten model-development feedback by identifying the failure mechanism that changed. Phase error and timing tails can guide minimization; density spikes provide a course-relative constraint; familiarity and typicality are soft targets or monitors. Optimization may use selected signals after task-specific stress testing, but no single number should become the reward.

pattern-IC band score

4-gram repetition / uniqueness

C4 loop collapse

self-similarity (higher looks better)

4-gram repetition / uniqueness

−1

0

1

Net direction vs. intact

Figure 3: Wrong-way incentives on the development panel. Common-pattern rewriting improves LM loss and its band score; loop collapse raises self-similarity. In both cases, 4-gram typicality falls. Bars show net direction from the intact chart; darker bars mark stronger corruption.

4

5

Discussion

AI Usage Statement

Typicality also needs context. A corruption often pushes a chart outside the same-course human band, but the reverse implication is invalid: an unusual chart may be deliberate and playable. Separate timing and feasibility checks help distinguish novelty from a concrete failure. When signals disagree, the disagreement is information rather than an error to hide with an average. The descriptive public-system profiles in Appendix C.3 illustrate two reusable reading habits. Similar clean-note fractions can hide different timing tails and global offsets, and a generator can match one difficulty while overshooting another. These examples are not additional metric validation; they show why diagnosis should follow failure mechanisms rather than one headline score.

6

AI assistants were used for experiment orchestration, drafting, and editing under author direction; the author checked all reported values.

Corpus Coverage and Held-Out Evaluation Protocol

A.1

Corpus coverage

The calibration corpus is summarized by course, star level, and tempo. Figure 4 shows its difficulty and tempo coverage.

A.2

Corpus and timing details

We use Taiko-style chart files and associated audio gathered from community resources available online. The training split contains 924 source rows that resolve to 813 unique song groups and 3,880 course charts. The test split contains 140 source rows that resolve to 120 song groups. The underlying charts and audio may be copyrighted, so they are not redistributed. The primary timing grid uses authored per-bar timestamps, meter, and tempo. For local timing diagnostics, every generated note is classified as within 6 ms, in the 6–12, 12–18, or above-18 ms range, or unmatched. We also compare adjacent intervals with tolerance max(6 ms, 2.5% × IOI). An audio-downbeat estimate provides a development-stage alternative: its period is within 2% of the metadata period on 39 of 40 development songs and differs by 2.39% on the remaining song. Held-out absolute offsets use the authored grid because it fixes bar phase directly.

Limitations

Controlled corruptions establish response to defined failures, not a universal ranking of chart quality. The present set does not cover every error, such as wrong-song pairing or removal of musically important notes. The responseto- music, long-form, and human-distance extensions were designed on the development panel and remain exploratory. The human bands come from one Taiko corpus. Timing also requires an external grid; without an authored map or a reliable audio estimate, the timing outputs are unavailable. Finally, timing clean rate has no reference-note recall term, so deleting off-grid notes can improve it. This is why the profile reads clean rate together with note rate and keeps feasibility checks separate.

7

A

Conclusion

ChartGenEval turns chart evaluation into separate, testable feedback signals. Seven output axes pass nine held-out controlled-corruption tests. On the development panel, the song-anchored phase estimate recovers injected shifts invisible to chart-only statistics. Those development tests also show that likelihood and self-similarity can move in the wrong direction under targeted failures, so neither should be used alone as a quality score. The broader lesson is simple: automatic evaluation is most useful when each signal names a failure, survives a targeted stress test, and keeps its role visible to the model developer.

A.3

Panel construction and integrity

Source rows are canonicalized by group_id and audio SHA-256; inconsistent aliases are rejected before panel indexing. The test split contains 120 canonical song groups in stored order. Groups 0–39 form the development panel, and groups 40–119 form the held-out panel. Before inspecting held-out results, we fixed the dataset revision, corruptions and doses, target measurements, controls, thresholds, calibration, and analysis plan.

A.4

Ethics Statement

Replicates, completeness, and inferential units

Every nonzero corruption cell and control condition uses five fixed base seeds. Replicates are averaged within chart and dose. A course is complete only when its intact value, every target dose and replicate, matched control, and required checks are available. A song enters a test with at least one complete course, and complete courses are averaged within song. Songs, rather than charts, courses, or replicates, are the inferential units for both response statistics. Deterministic no-op and missing cells do not enter the analysis.

The study uses community-transcribed charts and associated audio gathered from online sources as described in Section 3.3; these potentially copyrighted materials are not redistributed. The released materials exclude source charts, audio, personal data, and generators. Automated evaluation can lower the cost of content production, so high-volume deployment still requires platform policy and moderation.

5

(a) star level per course

(b) tempo distribution

Easy n=923

33

131

320

290

149

Normal n=924

11

55

91

204

231

205

127

Hard n=924

8

5

40

89

172

229

228

Oni n=924

3

3

4

7

21

median 170 BPM range 65–376

800

95

173

charts

600 153

274

180

400

164

200 Ura n=185

1

1

2

3

4

5

6

2

37

65

80

7

8

9

10

star level

0 100

200

300

400

beats per minute

Figure 4: Course difficulty and tempo coverage of the human reference corpus that defines the calibration bands. (a) Star-level counts per course across the 3,880 training charts: each course spans a broad, shifting difficulty range. (b) Tempo histogram (median 170 BPM, range 65–376).

A.5

A.7

Control checks and numerical invariance

The held-out execution produced 50,283 observation records from 333 charts. C1s retained 332 complete courses and C8 retained 330; every other effective pair retained 333. All 32 controls passed with maximum observed absolute difference zero. Nineteen deterministic no-op cells were excluded.

Identity re-evaluation applies to every retained measurement. A joint +1-second shift of notes, authored grid, and duration tests invariance to the time origin. A Don–Ka color bijection applies to timing and density, where color is irrelevant; it is not a grammar or form control. The 60 ms C2 translation is the matched control for chart-only density, grammar, and structure measurements. A required control failure prevents support for the affected target response. Each measurement uses a specified equivalence tolerance and timestamp-rounding rule. Chart-only interval arithmetic uses integer microseconds, three orders of magnitude finer than the smallest reported tolerance.

A.6

Sample counts and exclusions

B

Controlled-Corruption Specifications

Figure 5 visualizes the operational change made by each corruption and its prespecified target measurement. The 170-chart development matrix uses one realization per nonzero cell, producing 4,760 variants including the originals. C1s moves max(1, round(𝑛𝑝)) of a chart’s 𝑛 notes by ±60 ms. C3–C5 preserve note times, while C4 and C6 preserve or control density. Raw features and their derived scores are evaluated on the same variants; equivalent quantities are counted once. For C5, one LM chooses common-pattern rewrites and a separately fitted LM scores them. The perplexity reversal remains (9.44 → 6.04, −36%).

Intervals and decision rule

After orienting outputs so that larger is better, the first statistic is the within-chart Spearman correlation between the four severities (intact plus three doses) and the measurement; an all-constant response has zero sensitivity. The second statistic is the oriented maximum-dose value minus the matched control. Point estimates average song-level values. We use 10,000 song-cluster bootstrap draws: secondary analyses report nominal 95% intervals, and the ten primary rows use Bonferroni-adjusted 99.5% intervals. C6 assigns severities [0, 0.5, 0.5, 1] to the intact, 0.5×, 1.5×, and 2× conditions, so equal-magnitude low and high departures are tied; tied severities receive average ranks. A response is supported when at least 72 of 80 songs are evaluable, all controls pass, and both adjusted upper bounds are below zero. A nonnegative adjusted lower bound contradicts the response; other cases are inconclusive. The two C5 output rows were required to meet the rule together, although their algebraic equivalence leaves one effective axis. The C5 LM uses the full stored-order training split, trigram order 3, and add-𝛼 = 0.05.

6

Controlled corruption (strength)

Operational change

Prespecified response

C1 timing jitter (±10/20/30 ms) C1s sparse timing errors (0.5/1/2%, ±60 ms) C2 whole-chart shift (+15/30/60 ms) C3 note-type shuffle ( 𝑝=0.2/0.4/0.8) C4 loop collapse ( 𝑝=0.3/0.6/1) C5 common-pattern rewrite ( 𝑝=0.3/0.6/1) C6 note-rate scaling (×0.5/1.5/2) C7 burst insertion (2/4/8 clusters) C8 bar-order shuffle ( 𝑝=0.3/0.6/1)

all note times a few note times chart–song alignment local transitions local repetition pattern variety difficulty fit local overload phrase order

timing clean rate ↓ timing-error p99 ↑ grid-phase offset ↑ transition familiarity ↓ 4-gram typicality ↓ 4-gram typicality ↓ note-rate typicality ↓ density-spike limit ↓ 4-gram typicality ↓

Table 3: Controlled corruptions, strengths, and target readings. The edits target operational failures rather than claim to exhaust chart quality. C3–C5 preserve note times; C4 and C6 preserve or control density.

C1 dense timing jitter

C1s sparse 60 ms errors

C2 whole-chart shift

intact

intact

intact

edited

edited

edited

target: timing clean fraction

target: timing-error p99

target: chart-wide phase offset

C3 note-type shuffle

C4 loop collapse

intact

intact

ABCD

EFGH

intact

A

C

D

B

edited

edited

ABCD

ABCD

edited

A

B

D

C

target: transition familiarity

C5 common-pattern rewrite

target: 4-gram typicality (one effective axis)

target: 4-gram typicality

C6 note-rate scaling

C7 local burst insertion

C8 bar-order shuffle

intact

intact

intact

1

2

3

4

edited

edited

edited

3

1

4

2

target: note-rate typicality

target: density-spike limit

target: 4-gram typicality

Figure 5: Controlled-corruption atlas. Blue marks the intact chart and orange marks edited events. Each panel shows the property targeted by one corruption and names the target measurement. C1–C2 target timing, C3 targets transition familiarity, C4–C5 and C8 target repetition or form, and C6–C7 target difficulty fit and local load. These edits do not exhaust chart quality.

C

Metric Guide and Extended Results

C.1

Exploratory measurements

and rewards material that returns after an intervening phrase without rewarding adjacent copying. Stagnation– alienation marks active bars in long repeated runs or without a related bar within a 32-bar neighborhood. Density– energy response is the rank correlation between bar-level note density and mel energy. Run-head onset support asks whether the first note of a rhythmic group lies near strong spectral flux. Human-chart distance measures the mean

Five additional components were selected or revised on the 40-song development panel. They remain useful hypotheses, but the main paper does not treat them as held-outconfirmed targets. Reciprocity compares two-bar phrases 7

distance to the 20 nearest same-course training charts in a standardized 32-feature representation. Their development responses clarify their scope. Reciprocity reacts to loop collapse and bar shuffling; stagnation– alienation reacts most strongly to loop collapse. Run-head support detects whole-chart shifts, while density–energy response detects shuffled bars but is weak for short bursts. Human-chart distance detects many distributional departures but does not say whether an unusual chart is poor or creative. These outputs are therefore reported as exploratory monitors rather than core rewards.

C.2

Proxy diagnostics

The common-pattern rewrite is not an artifact of using the same LM for editing and scoring. When one LM selects the rewrite and a separately fitted LM scores it, perplexity still falls from 9.44 to 6.04 (−36%). A symmetric band transform of LM loss also fails to identify the maximum rewrite dose: its held-out change is +0.021 with a nominal 95% interval of [−0.034, 0.077]. The 4-gram axis supplies the missing reading. Under loop collapse, set-level SelfBLEU rises from 0.845 to 0.898, also recording the loss of diversity. Figure 6 connects each construct to its required inputs, representative outputs, and strongest present evidence. The evidence labels distinguish outputs with held-out corruption tests from components supported by development analyses. Figure 7 shows that targeted corruptions can affect multiple score families, motivating interpretation as a profile rather than as isolated interchangeable measurements.

8

Timing

Note sequence

Repetition & form

Are notes aligned to musical time?

Are local transitions familiar?

Does material return without stagnating?

A

clean fraction | residual tail chart-wide offset

held-out subset

LM fit | transition familiarity pattern typicality

held-out subset

B

A

4-gram typicality | reciprocity stagnation-alienation

B

held-out subset

Music response

Human-chart distance

Difficulty & limits

Does chart activity follow the song?

How far is the generated chart from same-course human charts?

Are chart demands typical for the requested course and within declared limits?

32-D nearest-neighbor gap

rate/load typicality overload and spike checks

density-energy response onset support

development

development

held-out subset

overload

4

-14.5

local density spike -3.6

4-gram repetition/uniqueness score

-2.7

-6.3

rare patterns

2

-8.7

note-rate score

3

1

short-interval score note-type switching

0

rhythm variety

−1

density variation local-pattern score

−2

-1.0

familiar transitions

−3

pattern irregularity local repetition

−4

large-note rate C1 jitter

C1s sparse

C2 shift

C3 types

C4 loop

C5 common

C6 rate

C7 burst

Change at strongest corruption (human SD)

Figure 6: Metric-to-construct guide. Each panel states an evaluation question, illustrates the relevant signal, lists representative outputs, and marks the strongest evidence. “Held-out subset” means that only the named outputs in Table 1 have held-out corruption evidence. Typicality describes departure from same-course human ranges.

C8 shuffle

Figure 7: How 14 nonredundant calibrated chart-only scores respond to each corruption at maximum strength on the development panel. Cells cover band-typicality scores and one-sided checks. Red means the score falls and blue that it rises; values are in standard deviations of the human reference charts, clipped at ±4. Black outlines mark the target rows selected on the development panel; they do not mark statistical significance. The matrix shows both target sensitivity and cross-family coupling.

9

C.3

Public-system profiles

We apply the profile to five public generators on the 40 development songs. Mapperatorinator v32 [17] and TaikoNation [2] produce Taiko note types. DDC’s onset pipeline [1], the public timing configuration of GenéLive! [3], and AutoOsu [18] enter only the applicable timing and audioresponse readings. We use their public inference configurations without training or fine-tuning on this corpus. These profiles illustrate interpretation; they are not held-out metric tests or a generator leaderboard. The two Taiko generators have similar median clean-note fractions, yet their relative-interval errors and global phase offsets differ. This is the same local-versus-global distinction exposed by the shift corruption. Difficulty conditioning reveals another pattern: Mapperatorinator overshoots human note rates most on easier courses, while its timing profile changes little by course. The useful conclusion is a course mismatch, not a general claim that the system improves at higher difficulty.

10

human reference (n=170)

DDC onset (n=40)

TaikoNation (n=40)

Mapperatorinator (n=170)

GenéLive! (n=200)

AutoOsu (n=40)

0.0

0.2

0.4

0.6

0.8

1.0

0.5

fraction of notes by timing-error range within 6 ms

6-12 ms

12-18 ms

1.0

within-6-ms fraction

>18 ms

unmatched

Figure 8: Timing profiles on the 40 development songs. The stacked bars separate locally clean notes from timing tails and unmatched notes; the right strip shows per-chart variation. Systems with similar mean clean fractions can have different tails, so one timing average does not identify the failure.

overload

density playability spike

human reference

1.00

1.00

1.00

0.75

1.00

1.00

Mapperatorinator

0.59

0.23

0.36

0.23

0.19

1.00

TaikoNation

0.58

1.00

1.00

0.00

0.01

0.08

Timing within 6 ms

Note rate

Short intervals

Familiar transitions

Local patterns

4-gram repetition / uniqueness

Figure 9: Separate development readings for human charts and the two generators that produce Taiko note types. Columns retain their own meanings; the matrix is a diagnostic profile, not a total score.

11

(a) NOTE-RATE RATIO

(b) CALIBRATED SCORES

(c) TIMING 1.0

3.5 Easy n=40

0.18

0.36

0.05

0.08

1.00

Normal n=40

0.11

0.14

0.01

0.04

1.00

Hard n=40

0.19

0.29

0.54

0.29

1.00

3.0

ratio

2.5

2.0

Oni n=40

1.5

1.00

1.00

0.76

1.00

1.00

fraction of notes

0.8

0.6

0.4

0.2 1.0

Ura n=10

1.00

1.00

note rate

short IOI

0.78

1.00

1.00

transition pattern familiarity fit

4-gram fit

0.0 E

N

H

O

course

U

E

N

H

O

U

course

Figure 10: Exploratory difficulty profile for Mapperatorinator on paired development charts. Easy through Oni contain 40 song–course pairs; Ura contains 10. Dots are medians and vertical lines span the 10th to 90th percentiles. Generated density approaches the human reference as course difficulty rises, while timing does not follow the same trend; this supports a course mismatch rather than general improvement at high difficulty.

12

D

Released Materials

[7] L.-C. Yang and A. Lerch, “On the evaluation of generative models in music,” Neural Computing and Applications, vol. 32, no. 9, pp. 4773–4784, 2020. [Online]. Available: https://doi.org/10.1007/ s00521-018-3849-7

Code, evaluation records, and plotting scripts are available at https://github.com/JacobLinCool/ ChartGenEval.

[8] H.-W. Dong, K. Chen, J. McAuley, and T. BergKirkpatrick, “MusPy: A toolkit for symbolic music generation,” in Proceedings of the 21st International Society for Music Information Retrieval Conference (ISMIR 2020), Montréal, Canada, 2020, pp. 101–108. [Online]. Available: https: //archives.ismir.net/ismir2020/paper/000187.pdf

References [1] C. Donahue, Z. C. Lipton, and J. McAuley, “Dance dance convolution,” in Proceedings of the 34th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, D. Precup and Y. W. Teh, Eds., vol. 70. Sydney, Australia: PMLR, Aug. 2017, pp. 1039–1048. [Online]. Available: https: //proceedings.mlr.press/v70/donahue17a.html

[9] S.-L. Wu and Y.-H. Yang, “The jazz transformer on the front line: Exploring the shortcomings of AI-composed music through quantitative measures,” in Proceedings of the 21st International Society for Music Information Retrieval Conference (ISMIR 2020), Montréal, Canada, 2020, pp. 142– 149. [Online]. Available: https://archives.ismir.net/ ismir2020/paper/000339.pdf

[2] E. Halina and M. Guzdial, “TaikoNation: Patterningfocused chart generation for rhythm action games,” in The 16th International Conference on the Foundations of Digital Games (FDG) 2021, ser. FDG ’21. New York, NY, USA: Association for Computing Machinery, Aug. 2021, pp. 1– 10. [Online]. Available: https://doi.org/10.1145/ 3472538.3472589

[10] M. Gover and O. Zewi, “Music translation: Generating piano arrangements in different playing levels,” in Proceedings of the 23rd International Society for Music Information Retrieval Conference (ISMIR 2022), Bengaluru, India, Dec. 2022, pp. 36–43. [Online]. Available: https://archives.ismir. net/ismir2022/paper/000003.pdf

[3] A. Takada, D. Yamazaki, Y. Yoshida, N. Ganbat, T. Shimotomai, N. Hamada, L. Liu, T. Yamamoto, and D. Sakurai, “Genélive! generating rhythm actions in Love Live!” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 4, pp. 5266–5275, Jun. 2023. [Online]. Available: https: //ojs.aaai.org/index.php/AAAI/article/view/25657

[11] J. Wang, W. Huang, and X. Li, “Mania archetype: Chart generation for rhythm action games with human factors,” in Human Factors in Virtual Environments and Game Design, vol. 137. AHFE Open Access, 2024, pp. 175–185. [Online]. Available: https://doi.org/10.54941/ahfe1005000

[4] J. Hanzen, E. Halina, and M. Guzdial, “Time-based chart partitioning: Improving local coherency in rhythm game chart generation,” Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment, vol. 21, no. 1, pp. 43–51, Nov. 2025. [Online]. Available: https: //ojs.aaai.org/index.php/AIIDE/article/view/36808

[12] M. Plasser, S. Peter, and G. Widmer, “Discrete diffusion probabilistic models for symbolic music generation,” in Proceedings of the ThirtySecond International Joint Conference on Artificial Intelligence (IJCAI-23). International Joint Conferences on Artificial Intelligence Organization, Aug. 2023, pp. 5842–5850. [Online]. Available: https://www.ijcai.org/proceedings/2023/648

[5] J. Li, M. Galley, C. Brockett, J. Gao, and B. Dolan, “A diversity-promoting objective function for neural conversation models,” in Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, K. Knight, A. Nenkova, and O. Rambow, Eds. San Diego, California: Association for Computational Linguistics, Jun. 2016, pp. 110–119. [Online]. Available: https: //aclanthology.org/N16-1014/

[13] J. Retkowski, J. Stępniak, and M. Modrzejewski, “Fréchet music distance: A metric for generative symbolic music evaluation,” arXiv preprint arXiv:2412.07948, Jan. 2025. [Online]. Available: https://arxiv.org/abs/2412.07948

[6] Y. Zhu, S. Lu, L. Zheng, J. Guo, W. Zhang, J. Wang, and Y. Yu, “Texygen: A benchmarking platform for text generation models,” in The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, ser. SIGIR ’18. New York, NY, USA: Association for Computing Machinery, Jun. 2018, pp. 1097–1100. [Online]. Available: https://doi.org/10.1145/3209978.3210080

[14] J. Opria, “STRUM: A spectral transcription and rhythm understanding model for end-to-end generation of playable rhythm-game charts,” arXiv preprint arXiv:2605.12135, May 2026. [Online]. Available: https://arxiv.org/abs/2605.12135

13

[15] A. Friberg and J. Sundberg, “Time discrimination in a monotonic, isochronous sequence,” The Journal of the Acoustical Society of America, vol. 98, no. 5, pp. 2524–2531, Nov. 1995. [Online]. Available: https://doi.org/10.1121/1.413218

[17] O. Schipper, “Mapperatorinator: An AI framework for generating and modding osu! beatmaps for all gamemodes from spectrogram inputs,” GitHub repository, May 2026, version 32.0.0; MIT-licensed software release. [Online]. Available: https://github.com/ OliBomby/Mapperatorinator/releases/tag/v32.0.0

[16] L. Theis, A. van den Oord, and M. Bethge, “A note on the evaluation of generative models,” in 4th International Conference on Learning Representations (ICLR 2016), Conference Track Proceedings, San Juan, Puerto Rico, May 2016. [Online]. Available: https://arxiv.org/abs/1511. 01844

[18] S. Lee and D. Jeong, “AutoOsu: Audio-aware action generation for rhythm games,” Late-Breaking Demo at ISMIR 2023, 2023, project page: https://issyun. github.io/autoosu/; public source repository: https: //github.com/issyun/AutoOsu. [Online]. Available: https://ismir2023program.ismir.net/lbd_319.html

14

Record · ID 366300 · SHA-256 402f1856b48ad93a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.