Conceptio › Archive › arXiv CS
arXiv CSopen access

MyoFlow: Anchor-Tied Rectified Flow for HD-sEMG Gesture Recognition Across Sessions and Subjects

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

MYOFLOW: ANCHOR-TIED RECTIFIED FLOW FOR HD-SEMG GESTURE RECOGNITION ACROSS SESSIONS AND SUBJECTS Chenhao Wu1 , Dingjie Peng1∗ , Satoshi Funabashi1 , Satoshi Konishi2 , Wuqiang Yang3 Hiroshi Onoda1 , Hironori Washizaki1 , Jiang Liu1

arXiv:2609.17194v1 [cs.LG] 15 Sep 2026

1

Waseda University, Japan

2

KDDI Research Inc., Japan

ABSTRACT High-density surface electromyography (HD-sEMG) gesture recognition supports prosthetic control, assistive robotics, and rehabilitation, but electrode re-donning and physiological variability cause distribution shifts that degrade accuracy across sessions and subjects. Generative HD-sEMG models primarily synthesize signals for augmentation; although diffusion models enhance representation learning, prediction still relies on a separate classifier. To tie learned dynamics to the decision rule, we propose MyoFlow, the first discriminative flow-matching framework for HD-sEMG recognition across sessions and subjects. It recasts classification as anchor-tied transport: a domain-conditioned rectified flow moves encoded windows toward gesture anchors that serve as transport targets and define the nearest-anchor decision geometry, enabling zero-shot prediction without an independent head. On the Hyser dataset, MyoFlow improves mean cross-session and cross-subject accuracy over the strongest diffusion-based baseline by 4.24% and 6.37%, respectively, and achieves 91.71% mean zero-shot accuracy and 97.39% mean few-shot accuracy across multiple days on the CEMHSEY dataset. Index Terms— HD-sEMG, gesture recognition, representation learning, flow matching, domain generalization. 1. INTRODUCTION High-density surface electromyography (HD-sEMG) records spatially resolved muscle activity through dense electrode arrays, providing a noninvasive interface for decoding motor intent [1, 2]. Models trained in a single recording domain, however, often generalize poorly to a new session or an unseen user [3]. Across sessions, electrode re-donning, changes in skin impedance, and movement variability alter the signal distribution [4]; across users, anatomical and neuromuscular differences introduce further shifts. To handle these shifts, previous methods combined handcrafted features with linear discriminant analysis or support vector machines [5, 6]. Although effective with limited data, their fixed features could not adapt to changes in the sensor ∗ Corresponding author

3

The University of Manchester, UK

interface or user physiology. Deep neural networks learn spatiotemporal features directly [7, 8], while domain adaptation and transfer learning further reduce source–target mismatch [9, 10]. Yet these methods focus on representations and decision boundaries rather than modeling the signal distribution, and they often require labeled target recordings for recalibration, which introduces extra user effort. Generative models offer a complementary route by modeling the distribution of HD-sEMG signals [11]. DiffHGR uses diffusion for HD-sEMG augmentation and feeds multiscale denoising features into the decoding branch of an auxiliary autoencoder [12]. At inference, however, prediction uses only the autoencoder’s encoder and a separately trained classifier; the diffusion is not used. This separation raises a question: Can the learned dynamics themselves make the class decision? Flow matching (FM) [13, 14] offers a potential route by learning a velocity field along predefined probability paths without trajectory simulation during training. However, its use in EMG remains limited to generation. EMGFlow [15] conditions noise-to-signal transport on a known gesture label, treating the label as an input rather than a prediction target and thus providing no direct classification rule. We therefore propose MyoFlow, to the best of our knowledge, the first FM framework for HD-sEMG classification across sessions and subjects. It recasts classification as subject- and sessionconditioned rectified-flow transport from encoded windows to a shared bank of learnable gesture anchors. These anchors serve as both transport endpoints and decision prototypes, allowing classification through nearest-anchor matching without a separate classifier head. Meanwhile, the shared geometry enables zero-shot recognition across domains without target-domain labels. Our main contributions are: ❶ We propose MyoFlow for HD-sEMG gesture recognition, formulating cross-session and cross-subject recognition as rectified transport in the representation space. ❷ We design an anchor-tied flow in which gesture anchors act as both transport targets and class prototypes, coupling representation learning to a nearest-anchor decision and making the learned class geometry identifiable. ❸ MyoFlow improves recognition on multiple datasets across sessions and subjects in zero- and few-shot settings.

t=0

t = 0.5

Pre -p

Session Code

roc

Cross Session

Train

Class Anchor 1

Subject Code

ess ing

Label

Train

Encoder W ind

ow s

t=1

Inference Initial Endpoints 𝑒!

Time

HD-sEMG Input

Transported Endpoint 𝑒"

Train

Conditional Flow Matching

Class Anchor 2

Cross Subject Class Anchor c

Label

Distance Calculation

Output

Fig. 1. The framework of the proposed MyoFlow for HD-sEMG gesture recognition across sessions and subjects.

2. METHOD

2.2. Conditional Rectified Flow

2.1. Problem Formulation Let x ∈ RM ×L denote an HD-sEMG window with M electrode channels and L temporal samples, labeled y ∈ {1, . . . , C} and tagged with the subject p and session d of its recording. Training uses only the labeled source set Dsrc . Cross-session recognition tests a source subject in a new session, and cross-subject recognition tests a subject absent from Dsrc . Evaluation is indexed by the number K of labeled target repetitions released for calibration: K = 0 denotes zero-shot; the K class repetitions form the support set and the remainder the query set for K > 0. The split is made at the repetition level, since windows from one repetition are strongly correlated without using query labels. The proposed MyoFlow realizes recognition as an initial-value problem in a latent space of endpoint dimension D = 256: e(0) = e0 = Eθ (x),

e(1) = eT ,

de(t) = vϕ (e(t), t, cp,d ) , dt

(1)

where Eθ denotes the initial endpoint encoder, and vϕ a velocity field conditioned on domain context cp,d . Following flow matching, t ∈ [0, 1] is the interpolation coordinate of a probability path: t = 0 is the encoded window and t = 1 the target distribution. The terminal endpoint eT = e(1) is thus where the class decision is read off. Fig. 1 provides an overview of our method. The anchor bank A = [a1 , . . . , aC ]⊤ ∈ RC×D holds one learnable anchor per gesture, shared across subjects and sessions, and supplies the targets of the path. Since one anchor bank supplies both the flow’s destinations and the classifier’s prototypes, representation learning and the decision rule are optimized on a single geometry. The encoder Eθ keeps time as the convolutional axis: a 1 × 1 stem mixes the M electrodes into feature channels, a stack of dilated depthwise temporal blocks widens the receptive field without striding, and global average pooling with a two-layer projection gives e0 ∈ RD .

For each labeled source window, MyoFlow pairs e0 with the anchor ay of its label, samples t ∼ U (0, 1), and defines the interpolated state and its constant target velocity as et = (1 − t)e0 + t sg[ay ],

ut = sg[ay ] − e0 ,

(2)

with sg[·] the stop-gradient. The field is fitted by 2

LFM = E(x,y,p,d)∼Dsrc , t ∥vϕ (et , t, cp,d ) − ut ∥2 .

(3)

As depicted in Fig. 1, the label selects the target anchor but never enters vϕ , so the field stays usable when the label is unknown; its squared-error optimum is the conditional mean E[ut |et , t, cp,d ], averaging the target velocities of paths meeting at the same state. Since ay is detached, the field learns to reach the anchors rather than drag them toward the data, while gradients through e0 shape the encoder and projector. The domain context cp,d is generated by a small MultiLayer Perceptron (MLP) from concatenated subject and session embeddings, enabling domain-specific transport under shared anchor geometry. At test time, an unseen session reuses the source-session embedding, while an unseen subject uses the mean training-subject embedding; neither requires target domain data. Zero-initializing the output layer of vϕ makes the initial transport map the identity. The FM loss requires no trajectory simulation, and the ordinary differential equation is solved only to obtain eT for the anchor-tied classifier in Sec. 2.3, using Euler integration with midpoint time sampling at four steps in training and eight in evaluation. 2.3. Anchor-Tied Classification The anchor bank is initialized as orthonormal√rows of a QR factorization, rescaled to a fixed radius r = D, so that all C(C − 1)/2 class pairs start equidistant. MyoFlow classifies the terminal endpoint eT by its distance to each anchor: ℓc (eT ) = −∥eT − ac ∥22 /τ,

ŷ = arg maxc ℓc (eT ), (4)

where ℓc is the logit of class c ∈ {1, . . . , C} and τ = D/sanchor sets the logit scale, with sanchor = 10. Expanding the squared distance yields a tied linear read-out with class weight 2ac /τ and bias −∥ac ∥22 /τ . Prediction is therefore equivalent to nearest-anchor classification. These weights are the flow’s own destinations; therefore, an invertible reparameterization of the endpoint space can no longer be absorbed into them. Only a global isometry leaves the objective unchanged, and pairwise anchor distances are invariant under it. The learned class geometry is therefore explicit rather than an artifact of the read-out. 2.4. Source-Only Learning No target-session window, statistic or label enters training. The source objective pairs flow matching with anchor supervision and auxiliary regularization: X Ltrain = λFM LFM + Lanchor + λr Lr , (5) r∈R

where Lanchor = λCE CE(ℓ(eT ), y) + λpull ∥eT − ay ∥22 /D, and R = {geo, rad, con} collects regularizers on class geometry, anchor norm, and clean–augmented consistency. The two leading terms divide the supervision: LFM constrains the path, Lanchor scores the arrival, and the latter is the only route by which the anchor bank is updated. Optimization runs in two source-only stages: shared pretraining over source sessions pooled across subjects, then, for cross-session evaluation, adaptation on the evaluated subject’s own source session. Leave-one-subject-out evaluation omits the second stage and excludes the held-out subject from pretraining, so that subject contributes no gradient at any point.

Table 1. Cross-session accuracy on Hyser (%). Method

0-rep (ZS) 1-rep (FS) 2-rep (FS) Avg.

ViT-MDHGR [7](2024) MoEMba [8](2025) DiffHGR [12](2026) MyoFlow (ours)

72.93 58.45 76.32 79.19

83.47 72.55 85.70 89.80

86.13 76.87 87.34 93.08

80.84 69.29 83.12 87.36

Table 2. Cross-subject accuracy on Hyser (%). Method

0-rep (ZS) 1-rep (FS) 2-rep (FS) Avg.

ViT-MDHGR [7](2024) MoEMba [8](2025) DiffHGR [12](2026) MyoFlow (ours)

60.90 46.76 63.37 64.12

76.71 66.35 78.01 86.55

80.74 72.36 80.30 90.11

72.78 61.82 73.89 80.26

3. EXPERIMENT SETUP We evaluate MyoFlow on the Hyser PR Dynamic [16] (20 subjects, two sessions separated by 3–25 days, 11 gestures, and 256 channels) and the CEMHSEY [17] (6 subjects, 11 consecutive days, 11 gestures, and 320 channels). Both datasets are sampled at 2048 Hz and segmented into nonoverlapping 50 ms windows. Cross-session evaluation uses Hyser Session 1 as the source and Session 2 as the target, while CEMHSEY uses Day 1 as the source and Days 2, 4, 6, 8, and 11 as targets. Cross-subject evaluation follows 20-fold leave-one-subject-out validation on Hyser Session 1, with no adaptation to the held-out subject. For K ∈ {0, 1, 2}, K complete target repetitions per class form the support set and the remainder form the query set. All experiments are conducted on a single NVIDIA RTX 5090 GPU.

2.5. Few-Shot Calibration 4. RESULTS AND ANALYSIS For K > 0, calibration acts on e0 rather than eT , since transport concentrates eT onto the anchors and leaves little withinclass variation to correct. From the labeled support alone we recompute BatchNorm statistics and initialize a linear calibration head gω from the class-wise support means µc , with wc = µc and bc = −∥µc ∥22 /2, which is exactly nearestprototype classification. The encoder, projector, and gω are then optimized on augmented support windows while the flow and anchor bank stay frozen; query repetitions are never seen. Let J (z) = CE(gω (z), y), φ(x) = e0 and φj (x) denote the pooled endpoint and the projection of the pre-pooling feature map at position j. The head gω is induced and supervised at both resolutions: i h λframe XL J (φj (x)) . (6) Lcal = E(x,y) J (φ(x)) + j=1 L The frame term draws L supervised signals from each window. And since the projector is non-linear, gω receives gradients from features that the pooled term never produces.

Overall comparison. As shown in Tables 1 and 2, MyoFlow outperforms every listed baseline across all Hyser settings. Relative to DiffHGR, the strongest baseline by average accuracy, it improves cross-session and cross-subject performance by 4.24 and 6.37 percentage points, respectively. For unseen subjects, MyoFlow exceeds the best 1- and 2-repetition baselines, DiffHGR and ViT-MDHGR, by 8.54 and 9.37 points. Table 3 further shows that MyoFlow achieves the highest zero-shot accuracy on every CEMHSEY target day, increasing the average from 87.52% to 91.71%. The benefit of 2-repetition calibration over its zero-shot result grows from 2.01 points on Day 2 to 8.44 points on Day 11. These findings reflect the complementary roles of the two stages: anchor-tied rectified transport establishes a shared decision geometry for zero-shot recognition, while initial-endpoint calibration uses target support to correct residual distribution shift. Temporal sensitivity. Fig. 2 shows that increasing the window length from 50 to 100 ms improves accuracy across

Table 3. Cross-session accuracy on CEMHSEY (%) for day gaps of 1–10 days. (Best performance; Second best performance) 0-rep (ZS)

Method

1-rep (FS)

2-rep (FS)

Day 2 Day 4 Day 6 Day 8 Day 11 Avg. Day 2 Day 4 Day 6 Day 8 Day 11 Avg. Day 2 Day 4 Day 6 Day 8 Day 11 Avg.

Accuracy (%)

ViT-MDHGR [7] (2024) MoEMba [8] (2025) DiffHGR [12] (2026) MyoFlow (ours)

100 90 80 70 60 50

90.57 78.88 92.69 96.06

86.63 72.60 85.56 90.84

85.97 70.69 87.90 91.87

84.05 68.87 84.80 89.66

83.00 68.22 86.67 90.11

86.04 71.85 87.52 91.71

0-rep (ZS)

1-rep (FS)

2-rep (FS)

****

***

***

50

ns

100 150

Window (ms)

50

ns

100 150

Window (ms)

50

96.38 89.34 96.46 97.41

True Positive Rate (TPR)

Hyser (macro AUC = 0.963)

1.0

0.8

0.4 0.2 0.0 −4 10

10−3

10−2

10−1

False Positive Rate (FPR)

94.87 87.92 97.14 97.42

95.25 88.66 96.52 96.95

97.38 91.31 97.40 98.07

95.28 91.21 96.42 97.14

96.26 92.42 97.69 97.37

96.57 93.47 97.39 98.01

96.63 92.94 97.89 98.55

96.42 92.27 97.36 97.83

anchor

100 150

Fig. 4. Visualization of the rectified-flow transport that carries each initial endpoint to its class anchor.

CEMHSEY (macro AUC = 0.992)

0.6 Class 1 (AUC = 0.955) Class 2 (AUC = 0.965) Class 3 (AUC = 0.967) Class 4 (AUC = 0.979) Class 5 (AUC = 0.952) Class 6 (AUC = 0.960) Class 7 (AUC = 0.977) Class 8 (AUC = 0.969) Class 9 (AUC = 0.963) Class 10 (AUC = 0.950) Class 11 (AUC = 0.958)

96.62 89.94 97.73 98.22

Window (ms)

0.8

0.6

94.75 88.48 96.81 96.77

endpoint

ns

Fig. 2. Temporal-window sensitivity for cross-session recognition on Hyser, with pairwise significance tests.

1.0

93.62 87.63 94.44 94.92

Class 1 (AUC = 0.995) Class 2 (AUC = 0.990) Class 3 (AUC = 0.995) Class 4 (AUC = 0.991) Class 5 (AUC = 0.995) Class 6 (AUC = 0.992) Class 7 (AUC = 0.985) Class 8 (AUC = 0.988) Class 9 (AUC = 0.993) Class 10 (AUC = 0.997) Class 11 (AUC = 0.989)

0.4 0.2

0.0 100 10−4

10−3

10−2

10−1

False Positive Rate (FPR)

100

Fig. 3. Zero-shot Receiver Operating Characteristic (ROC) curves of MyoFlow on both datasets. (The FPR is log-scaled)

ination under both cross-session and cross-day distribution shifts, complementing the multiclass accuracy results. Interpretability of learned transport. Fig. 4 visualizes the endpoint trajectories projected onto the two-dimensional principal subspace of the anchor bank. As flow progresses, the endpoints move toward and concentrate around their class anchors, revealing how rectified transport organizes the latent space into a decision-aligned geometry. Because prediction uses the same anchors, each trajectory directly connects representation dynamics to the final decision, providing a geometric, sample-level interpretation of recognition. 5. CONCLUSION

all settings, as confirmed by paired Wilcoxon signed-rank tests, whereas extending it to 150 ms yields no significant further gain. These results reveal an accuracy–latency trade-off: compared with 100 ms, the 50 ms setting halves acquisition time while retaining strong recognition accuracy, supporting deployment in latency-sensitive settings. Classifier discrimination. We construct class-wise onevs-rest ROC curves by pooling target windows within each dataset and sweeping thresholds over the per-window, zeroshot, anchor-derived probabilities. Fig. 3 reports macroAUCs of 0.963 on Hyser and 0.992 on CEMHSEY, with every class-wise AUC exceeding 0.950. These results show that the anchor-tied readout maintains reliable class-wise discrim-

We introduced MyoFlow, the first discriminative flow matching framework for cross-session and cross-subject HD-sEMG gesture recognition. Its anchor-tied design embeds transport in the decision rule: a domain-conditioned rectified flow moves encoded windows toward gesture anchors serving as transport endpoints and prototypes, enabling nearest-anchor prediction without an independent head. This geometry supports zero-shot transfer, while few-shot calibration adapts the variation-preserving initial endpoint to residual distribution shifts. Experiments on Hyser and CEMHSEY show strong recognition under session and subject shifts, with the largest calibration gains at longer day gaps. Together, these results show that learned transport can define the class decision rather than remain auxiliary to a separate classifier.

Acknowledgement This work was supported by the Waseda University Grant for Special Research Projects (Project No. BARH02612801), the Waseda Research Institute for Science and Engineering, and JST BOOST under Grant Number JPMJBS2429.

References [1] Wei Li, Ping Shi, and Hongliu Yu, “Gesture recognition using surface electromyography and deep learning for prostheses hand: State-of-the-art, challenges, and future,” Frontiers in Neuroscience, vol. 15, 2021. [2] Jehan Yang, Kent Shibata, Douglas Weber, and Zackory Erickson, “High-density electromyography for effective gesture-based control of physically assistive mobile manipulators,” npj Robotics, vol. 3, no. 1, 2025. [3] Wenhui Cui, Christopher M Sandino, Hadi Pouransari, Ran Liu, Juri Minxha, Ellen Zippi, Erdrin Azemi, and Behrooz Mahasseni, “Embridge: Enhancing gesture generalization from emg signals through cross-modal representation learning,” in International Conference on Learning Representations, 2026, vol. 2026, pp. 149668– 149684. [4] Fang Qiu, Xiaodong Liu, and Xinming Ye, “Influence of electrode placement on the recognition of different gesture categories using high-density sEMG,” Frontiers in Neuroscience, vol. 19, 2025. [5] Bernard Hudgins, Philip Parker, and Robert N. Scott, “A new strategy for multifunction myoelectric control,” IEEE Transactions on Biomedical Engineering, vol. 40, no. 1, pp. 82–94, 1993. [6] Kevin Englehart and Bernard Hudgins, “A robust, realtime control scheme for multifunction myoelectric control,” IEEE Transactions on Biomedical Engineering, vol. 50, no. 7, pp. 848–854, 2003. [7] Qin Hu, Golara Ahmadi Azar, Alyson Fletcher, Sundeep Rangan, and S Farokh Atashzar, “Vit-mdhgr: Cross-day reliability and agility in dynamic hand gesture prediction via hd-semg signal decoding,” IEEE Journal of Selected Topics in Signal Processing, vol. 18, no. 3, pp. 419–430, 2024. [8] Mehran Shabanpour, Kasra Laamerad, Sadaf Khademi, and Arash Mohammadi, “Moemba: a mamba-based mixture of experts for high-density emg-based hand gesture recognition,” in 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC). IEEE, 2025, pp. 1–6.

[9] Md. Rabiul Islam, Daniel Massicotte, Philippe Massicotte, and Wei-Ping Zhu, “Surface EMG-based intersession/intersubject gesture recognition by leveraging lightweight All-ConvNet and transfer learning,” IEEE Transactions on Instrumentation and Measurement, vol. 73, pp. 1–16, 2024. [10] Jianfeng Li, Xinyu Jiang, Jiahao Fan, Yanjuan Geng, Fumin Jia, and Chenyun Dai, “Deep end-to-end transfer learning for robust inter-subject and inter-day hand gesture recognition using surface EMG,” Biomedical Signal Processing and Control, vol. 100, pp. 106892, 2025. [11] Nana Wang, Suli Wang, Gen Li, Pengfei Ren, and Hao Su, “New synthetic goldmine: Hand joint angle-driven emg data generation framework for microgesture recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence, 2026, vol. 40, pp. 17787–17795. [12] Kejia Su, Bo Wan, Jiayang Huang, Zhi-Qiang Zhang, Pengfei Yang, and Quan Wang, “Diffusion-based learning for cross day hand gesture recognition using HDsEMG signals,” Biomedical Signal Processing and Control, vol. 118, pp. 109716, 2026. [13] Yaron Lipman, Ricky T. Q. Chen, Heli Ben-Hamu, Maximilian Nickel, and Matthew Le, “Flow matching for generative modeling,” in International Conference on Learning Representations, 2023. [14] Xingchao Liu, Chengyue Gong, and Qiang Liu, “Flow straight and fast: Learning to generate and transfer data with rectified flow,” in International Conference on Learning Representations, 2023. [15] Boxuan Jiang, Chenyun Dai, and Can Han, “Emgflow: Robust and efficient surface electromyography synthesis via flow matching,” arXiv preprint arXiv:2604.13685, 2026. [16] Xinyu Jiang, Xiangyu Liu, Jiahao Fan, Xinming Ye, Chenyun Dai, Edward A Clancy, Metin Akay, and Wei Chen, “Open access dataset, toolbox and benchmark processing results of high-density surface electromyogram recordings,” IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 29, pp. 1035– 1046, 2021. [17] Shutian Yang, Chen Chen, Dongxuan Li, and Xiangyang Zhu, “A consecutive multi-day high-density surface electromyography dataset comprising 7 grasps and 11 gestures,” Scientific Data, vol. 12, no. 1, pp. 1420, 2025.

Record · ID 919395 · SHA-256 4d05e33b7d1039c5
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.