Conceptio › Archive › arXiv CS
arXiv CSopen access

FROD: Feature Matching Residual Denoising Oracle Bone Decipher

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

FROD: Feature Matching Residual Denoising Oracle Bone Decipher

arXiv:2609.17227v1 [cs.CV] 15 Sep 2026

Yanbin Hou1 , Biao Xiong1 , Guojun Xu1 , Jianwen Xiang1 , Cheng Tan1 , Yanchao Yang1 , and Junwei Zhou1(B) School of Artificial Intelligence, Wuhan University of Technology, Wuhan 430070, China [email protected], [email protected], [email protected], [email protected], [email protected], [email protected], [email protected]

Abstract. Oracle bone script (OBS), one of the earliest Chinese writing systems, plays an important role in the study of Chinese etymology. Traditional decipherment relies heavily on domain experts who analyze characters through semantic context and structural evolution. To assist this labor-intensive process, we formulate OBS decipherment assistance as a cross-era image translation task and propose FROD (Feature Matching Residual Denoising Oracle Bone Decipher). Although many OBS characters differ substantially from their modern counterparts, they often preserve local topological invariants at the radical level. During training, FROD leverages fast feature matching to provide gated segmentation supervision: paired samples with sufficient matches are processed patch-wise to align fine-grained radicals, whereas low-similarity pairs are trained holistically to avoid mismatched artifacts. In addition, a Residual Denoising Diffusion Model (RDDM) jointly estimates noise and residual signals, thereby reducing the positional drift and stroke disorder commonly observed in standard diffusion models. Finally, a multi-stage font stylization refinement network refines the generated images by eliminating edge noise and stabilizing stroke structures. On our augmented character-disjoint dataset, FROD achieves higher Top-1 recognition accuracy than the evaluated baselines, with a 3.8% absolute gain over OBSD. Keywords: oracle bone script, image translation, diffusion model, feature matching

1

Introduction

Oracle bone script is one of the earliest forms of pictographic writing and was used during the Shang Dynasty. The study of oracle bone script (OBS) is fundamental to research on Chinese etymology and ancient history. However, among the approximately 4,500 unique characters discovered so far, only about 1,500 have been reliably deciphered. The decipherment process remains laborintensive, requiring substantial domain knowledge and associative reasoning to map ancient symbols to their modern counterparts.

2

Y. Hou et al.

Although end-to-end decipherment ultimately depends on contextual and linguistic evidence, generating visually recognizable modern Chinese character candidates from OBS rubbings can greatly accelerate the research workflow. Recent studies have therefore formulated this problem as an image-to-image translation task. In particular, OBSD [1] employs a diffusion model with Local Structural Sampling (LSS) to align fine-grained patches between OBS and modern characters. However, OBS characters often exhibit substantial structural variation. Forcing highly abstract or weakly aligned pairs into naive patch-wise mappings can lead to severe positional drift, mismatched artifacts, and stroke disorder. In such cases, blind segmentation becomes unreliable because explicit structural correspondences are weak. To address these challenges, we propose FROD, an image translation framework for OBS decipherment assistance. Compared with OBSD, our approach offers three key improvements. First, rather than blindly segmenting all training pairs, we use the fast feature-matching model LightGlue [2] to provide gated segmentation supervision. During training, paired OBS and modern characters with sufficient topological overlap, such as shared radicals, are processed patchwise to align local features, whereas highly variable pairs are trained holistically to preserve global structure. Second, we adopt the Residual Denoising Diffusion Model (RDDM) [3] in place of a standard diffusion model. By modeling target residuals and generated noise as separate components, RDDM alleviates the positional drift and structural bias that often arise in conventional noise-prediction frameworks. Third, we apply a multi-stage font stylization refinement network [4] to refine the synthesized outputs, suppress edge noise, and produce standardized modern glyphs. Our main contributions are as follows:

– We introduce a feature-matching-gated training strategy for OBS image translation. By assigning paired training samples to patch-wise or holistic supervision according to structural similarity, our method avoids detrimental misalignments while capturing local radical correspondences when appropriate. – We adapt a Residual Denoising Diffusion Model to the OBS domain, which explicitly models residual signals to better capture stroke positional patterns and preserve structural fidelity compared to standard diffusion baselines. – We integrate a font stylization refinement module to refine the predicted glyphs. Extensive experiments show that the full pipeline achieves the best OCR accuracy among the evaluated baselines on our augmented characterdisjoint dataset and improves Top-1 OCR accuracy over OBSD. By producing reliable modern character candidates, FROD helps bridge computer vision and archaeology and provides a practical assistive tool for OBS decipherment.

FROD

2

3

Related Work

OBS research has advanced substantially with the release of several digitized datasets. Representative collections include Oracle-20K [5], OBC306 [6], and HWOBC [7]. However, these early datasets focus primarily on isolated character recognition and do not provide structural evolutionary mappings between OBS and modern Chinese characters. More recent resources, such as EVOBC [8], HUST-OBC [9], and OBC-V [10], address this limitation by providing moderncharacter correspondences and expanded class coverage, thereby laying the foundation for translation-based approaches. For image-to-image translation, conditional Generative Adversarial Networks (GANs) such as Pix2Pix [11], CycleGAN [12], and DRIT++ [13], as well as diffusion models such as Palette [14] and BBDM [15], have shown strong performance. However, they often struggle to align highly deformed cross-era topologies without explicit structural guidance. Within OBS studies, most prior work has focused on character recognition using computer vision or natural language processing techniques, whereas AIassisted interpretation of undeciphered characters remains limited. Zhang et al. [16] proposed a case-based reasoning method that retrieves structurally similar cases from adjacent writing systems to assist expert decipherment. Chang et al. [17] introduced a cascaded GAN framework that models intermediate stages of Chinese-character evolution from OBS to modern forms. Guan et al. [1] introduced a diffusion-based approach for detailed OBS-modern character alignment, thereby providing clearer evolutionary links. Nevertheless, existing methods often struggle to preserve fine-grained structural differences between OBS and modern character images, which leads to positional drift, missing strokes, and blurred outputs. In addition, few pipelines include explicit font standardization after generation, which further limits the legibility of translated characters for downstream recognition.

3

Feature Matching and Segmentation

Due to the highly variable structure of OBS characters, establishing valid correlations with modern Chinese character images is essential. During paired training, blindly applying patch-wise supervision to highly abstract OBS-modern pairs often produces disordered or noisy strokes because blind segmentation forces alignments where explicit structural correspondences do not exist. To address this issue, we introduce a training-time gating mechanism driven by fast feature matching. The number of high-confidence matched keypoints serves as a quantitative criterion for determining whether a paired training sample should receive patch-wise local supervision or holistic supervision. 3.1

Feature Matching and Gating Mechanism

We adopt LightGlue (LG) [2] to extract and match keypoints due to its efficiency on low-complexity images like skeletonized OBS strokes. LG utilizes self-

4

Y. Hou et al.

Fig. 1. Training-time gating mechanism for OBS (left) and modern characters (right). Top: dense keypoint matches (nmatch > γ) trigger patch-wise local segmentation for radical alignment. Bottom: sparse matches (nmatch ≤ γ) lead to holistic processing to preserve global character topology, where γ = 15.

and cross-attention layers enhanced by Rotary Positional Embeddings to capture spatial context. Based on these contextualized features, a lightweight head predicts point correspondences. Let xIi denote the descriptor of the i-th keypoint from the modern character image and xSj denote the descriptor of the j-th keypoint from the paired OBS image, where M and N are the corresponding numbers of detected keypoints. The similarity score matrix S ∈ RM ×N is computed as Sij = (Wmatch xIi )⊤ (Wmatch xSj ).

(1)

With point-wise matchability confidences βiI , βjS ∈ [0, 1], the soft assignment matrix O is formulated as Oij = βiI · βjS · Softmaxk∈I (Skj )i · Softmaxk∈S (Sik )j .

(2)

A match is considered valid if it is mutually maximal and its score Oij strictly exceeds a confidence threshold τ . The number of valid matches is not treated as a semantic equivalence measure; instead, it serves as a practical proxy for the reliability of local structural correspondences between paired glyph images. Gating Threshold (γ): During training, we count the total number of valid matches for each paired OBS-modern image sample. We set γ = 15 as a conservative empirical threshold to select pairs with sufficiently dense local correspondences for patch-wise supervision. If the match count exceeds this threshold, the pair exhibits sufficient structural correlation and is routed to the patch-wise segmentation module. Otherwise, it bypasses segmentation and is trained holistically to preserve global topology. As shown in Fig. 1, the upper examples exceed the threshold γ, triggering segmentation, while the lower ones do not. 3.2

Image Segmentation Method

For paired training samples that satisfy the gating condition, we adopt an LSS strategy. Specifically, the resized 100×100 input image I ∈ R100×100×C is divided

FROD

5

Fig. 2. Steps of the residual denoising diffusion model.

into D = 8 overlapping local regions of size p × p (with p = 64) using a sliding window. To obtain exactly eight local regions while maintaining sufficient coverage of the character strokes, we use a grid sampling strategy with appropriate strides. In implementation, each local region is represented on a fixed 100 × 100 canvas before being fed into the diffusion network, while holistic samples also use the original 100 × 100 image. Thus, patch-wise and holistic training share the same network input resolution; p only defines the local support region used for segmentation and blending. (d) These local patch representations, denoted as I0 for the modern character (d) and I˜ for the corresponding OBS character, are processed independently during training. A Gaussian blending mask Pd is generated for each patch position, which is later utilized during the reverse generation process to seamlessly blend overlapping regions and prevent visible seam artifacts.

4

Residual Denoising Diffusion Method

Traditional conditional diffusion models typically estimate either the target image or the injected noise. Although methods such as I2 SB [18] model bridges between paired image domains, they do not explicitly decouple deterministic domain residuals from stochastic noise. For OBS decipherment, where the structural gap between paired glyphs can be large, we adapt the RDDM framework of Liu et al. [3] to OBS image translation. By explicitly treating the target residual and generated noise as independent components, this dual-estimation framework preserves the OBS structural condition more faithfully. The overall process is illustrated in Fig. 2. 4.1

Residual Denoising Formulation

Following RDDM [3], we redefine the diffusion process to transition from the target modern character image I0 to a state heavily dependent on the conditional ˜ We define the domain residual as Ires = I˜ − I0 . OBS image I. Forward Diffusion Process: The forward process gradually adds noise ˜ The marginal distribution at step t is formulated and shifts the mean toward I.

6

Y. Hou et al.

Algorithm 1 Patch-wise Reverse Sampling with Gaussian Blending ˜ conditional networks ϵθ (·, ·, t) and Ires,θ (·, ·, t); Gaussian blending Input: OBS image I; masks Pd Output: Preliminary translated modern character I0 1: Sample initial noise zT ∼ N (0, I) 2: Initialize IT ← I˜ + β̄T zT 3: Set stability constant c ← 10−8 4: for t = T downto 1 do 5: Initialize accumulators Σt ← 0, Φt ← 0, and M ← 0 6: for d = 1 to 8 do (d) 7: Extract local region d from It and embed it into a fixed-size canvas It (d) 8: Extract local region d from I˜ and embed it into a fixed-size canvas I˜ (d) (d) 9: Predict noise ϵ̂t ← ϵθ (It , I˜(d) , t) (d) (d) 10: Predict residual Iˆres,t ← Ires,θ (It , I˜(d) , t) (d) (d) 11: Crop the valid local support from ϵ̂t and Iˆres,t , and project it back to local region d (d) 12: Σt ← Σt + Projd (ϵ̂t ) ⊙ Pd (d) 13: Φt ← Φt + Projd (Iˆres,t ) ⊙ Pd 14: M ← M + Pd 15: end for 16: ϵ̄t ← Σt /(M + c) 17: I¯res,t ← Φt /(M + c) 18: Sample step noise zt ∼ N (0, I) if t > 1, else 0 ¯ 19: res_term ← (ᾱ t − ᾱt−1 ) · Ires,t q  2 20: noise_term ← β̄t − β̄t−1 − σt2 · ϵ̄t 21: It−1 ← It − res_term − noise_term + σt · zt 22: end for 23: return I0

consistently using schedule parameters ᾱt and β̄t : q(It | I0 , Ires ) = N (It ; I0 + ᾱt Ires , β̄t2 I),

(3)

where I denotes the identity matrix. The parameter ᾱt ∈ [0, 1] monotonically increases to 1 at t = T , and β̄t represents the noise scale. Using the reparameterization trick, we can express It as It = I0 + ᾱt Ires + β̄t ϵ,

ϵ ∼ N (0, I).

(4)

Notice that when t = T , ᾱT ≈ 1, leading to the boundary condition IT ≈ I˜+ β̄T ϵ. This ensures that the reverse process initiates from a noisy version of the OBS image rather than pure Gaussian noise. Reverse Generation Process: To reverse the process from IT to I0 , the model employs two coupled networks to predict the residual and noise simulta˜ t) and ϵθ (It , I, ˜ t). The target image is estimated at each step neously: Ires,θ (It , I, as I0,θ = It − ᾱt Ires,θ − β̄t ϵθ . (5)

FROD

7

˜ is parameterized as N (It−1 ; µθ , σt2 I). The reverse transition step pθ (It−1 | It , I) Deriving from the posterior distribution q(It−1 | It , I0 , Ires ) and substituting our estimates, the mean µθ is computed. The final sampling step becomes   q 2 2 (6) It−1 = It − (ᾱt − ᾱt−1 )Ires,θ − β̄t − β̄t−1 − σt ϵθ + σt z, where z ∼ N (0, I), and σt2 is the reverse step variance controlled by a stochasticity parameter η ∈ [0, 1]. Following the RDDM sampling schedule, σt is chosen 2 such that β̄t−1 − σt2 ≥ 0, which keeps the square-root term well-defined. The network is optimized using a weighted combination of the residual loss and noise loss: L(θ) = Lres (θ) + λLϵ (θ),   2 ˜ Ires − Ires,θ (It , I, t) , Lres (θ) = Et,I0 ,I,ϵ ˜ 2   2 ˜ t) Lϵ (θ) = Et,I0 ,I,ϵ ϵ − ϵθ (It , I, . ˜

(7) (8) (9)

2

4.2

Patch-wise Training Loss Adaptation

As discussed in Sect. 3.2, the paired OBS and target images are segmented into (d) D local patches I˜(d) and I0 under the LSS strategy. The training objective averages prediction errors across patches and time steps. Accordingly, the expected loss functions are computed in a patch-wise manner: L̂(θ) = L̂res (θ) + λL̂ϵ (θ),   2 (d) ˜(d)  (d) L̂res (θ) = Et,d,I0 ,I,ϵ Ires − Ires,θ It , I , t , ˜ 2    2 (d) ϵ(d) − ϵθ It , I˜(d) , t . L̂ϵ (θ) = Et,d,I0 ,I,ϵ ˜

(10) (11) (12)

2

The goal is to minimize the difference between predicted and ground truth patch (d) residuals Ires and noise ϵ(d) . Averaging the prediction errors across patches helps the model learn the localized distribution of structural correspondences, preventing overfitting to individual patches. Here, each patch residual is defined as (d) (d) Ires = I˜(d) − I0 , consistent with the global residual definition. For pairs that do not meet the gating threshold, no spatial decomposition is applied and D is set to 1. This allows the same loss formulation to cover both patch-wise and holistic training samples. 4.3

Inference Process

At inference time, since paired modern targets are unavailable, the LightGluebased gate is not used. FROD adopts a fixed patch-wise reverse sampling strategy

8

Y. Hou et al.

Fig. 3. Steps of font stylization refinement.

for every OBS input and reconstructs the final output using Gaussian blending. Specifically, the input OBS image I˜ is decomposed into eight overlapping local regions, and each local region is translated at the same 100 × 100 network resolution as holistic processing. This inference-time patch decomposition does not impose additional paired local correspondence supervision; it serves as an input decomposition and blending strategy for recovering local strokes while preserving global consistency. Algorithm 1 summarizes the patch-wise reverse sampling procedure. During training, gated samples use the patch-wise losses in Sect. 3.2, whereas low-match samples use the same loss with D = 1 as holistic supervision.

5

Font Stylization Refinement

Although RDDM substantially improves the recovery of modern character topology from oracle bone script, the preliminary outputs often retain raw brush textures, structural fragmentation, and artifacts inherited from the ancient domain. To bridge the large domain gap between these coarse outputs and standard printed modern fonts, which is crucial for reliable downstream OCR rather than mere visual refinement, we introduce a multi-stage font stylization refinement network. This stage preserves the core character structure while regularizing stroke widths and removing raster artifacts. Inspired by few-shot font generation methods, we adopt the MSD-Font architecture proposed in [4]. The font stylization refinement module maps coarse translated characters to clean, stylized modern fonts and thereby unifies visual representations for fair evaluation. Although the modern font corpus covers the evaluated modern character categories, it contains only standard printed glyphs and no OBS images or OBS-modern paired samples. The font stylization refinement module is there-

FROD

9

fore trained independently as a modern-font structural and stylistic prior, rather than as OBS-to-modern paired supervision. During inference, the style image only provides font-level appearance guidance and is not the paired target glyph of the input OBS sample. This module serves as a post-generation standardization stage that regularizes the coarse RDDM output into a clean modern glyph form. The overall MSD-Font architecture is shown in Fig. 3. We use the preliminary translation from the RDDM stage as the source image Is . The source image is encoded by a Vector Quantized VAE (VQ-VAE) encoder into the latent feature z̃0s . In parallel, a style image Ig is processed by the style encoder to produce the style condition yi . The style image provides font-level appearance guidance and is not used as a paired target glyph for the test OBS input. During the forward diffusion process, Gaussian noise is added to z̃0s to obtain the intermediate latent state z̃t2 . During reverse diffusion, MSD-Font employs a style-conditioned prediction network z̃ (r,i) (z̃t , t, yi ) to progressively transform the source latent toward a clean stylized result, and the VQ-VAE decoder finally outputs the generated modern glyph Ir . The reverse process consists of three distinct stages: glyph construction, font transformation, and refinement. Starting from the noisy latent state associated with the source image, the model first reconstructs the coarse content structure of Is during the glyph construction stage to anchor the basic topology. This stage simplifies to a single-step forward transition at an intermediate timestep t2 : p p (13) z̃t2 = ᾱt2 z̃0s + 1 − ᾱt2 ϵ0 , where ᾱt2 is the predefined cumulative diffusion coefficient at time t2 , and ϵ0 ∼ N (0, I). After obtaining the first-stage latent noise map z̃t2 , the font transformation stage progressively transforms it into the intermediate stylized latent z̃t1 using a style-conditioned network z̃ (r,1) (z̃t , t, y1 ). Finally, the font refinement stage uses a secondary conditional network z̃ (r,2) (z̃t , t, y2 ) to refine z̃t1 , repairing local stroke intersections and yielding the final latent representation for the generated character. The VQ-VAE decoder then produces the final clear, stylized modern Chinese character image Ir .

6

Experimental Results and Analysis

6.1

Experimental Settings

During training, the diffusion model is optimized with an initial learning rate of 2 × 10−4 . We maintain an Exponential Moving Average (EMA) of model parameters with a decay rate of 0.995 to improve optimization stability. Training is conducted for 350 epochs with a batch size of 16. All network inputs are resized to 100×100 pixels for both patch-wise and holistic processing. For gated samples, we use eight overlapping local regions with p = 64 as the segmentation support, and the segmentation threshold is set to γ = 15.

10

6.2

Y. Hou et al.

Datasets and Evaluation Metrics

We use OBC-V [10] as the base dataset and augment it with supplementary OBS images from EVOBC [8] and HUST-OBC [9]. This augmentation is performed by adding real OBS samples from external datasets rather than synthesizing new glyph images. After duplicate removal, the augmented dataset contains 74,219 OBS images mapped to 1,590 interpreted modern Chinese character classes. To prevent class-level leakage and ensure rigorous evaluation, we implement a character-disjoint train-test split while maintaining an approximately 9:1 split ratio. Specifically, all OBS samples belonging to the same modern Chinese character class are assigned exclusively to either the training set or the test set, and no modern character class appears in both sets. We further remove duplicate OBS instances before splitting and ensure that instances extracted from the same OBS source image are not shared across splits. We evaluate the framework using both image quality metrics and recognition accuracy. Image Quality Metrics: We report standard metrics, including FID, RMSE, SSIM, and LPIPS. The generated characters are compared directly with the corresponding standard printed modern Chinese fonts used as ground-truth references. Recognition Accuracy (Top-k): To evaluate isolated-glyph automatic decipherment objectively, we deploy an OCR model. Specifically, we train a ResNet-50 classifier on a standard corpus of printed modern Chinese characters covering all 1,590 evaluated classes. The OCR classifier is trained on standard printed modern Chinese fonts and is used solely as an automatic recognizer for generated modern glyphs. The generated translation outputs are fed into this classifier to compute Top-1, Top-5, Top-10, Top-100, and Top-500 accuracies. For the OBSD [1] baseline, we use the official implementation and retain its LSS-based initial decipherment and zero-shot refinement stages, so OBSD is evaluated as its complete published pipeline. 6.3

Comparison and Analysis of Results

Table 1 reports the quantitative image translation results. FROD achieves the best performance across all image-quality metrics. GAN-based methods, including Pix2Pix, CycleGAN, and DRIT++, perform poorly because they have limited capacity to model the severe structural deformations between OBS and modern characters. Diffusion-based methods generalize better, and FROD further improves over the strongest baseline, OBSD, by learning more robust structural priors through RDDM. Table 2 summarizes the OCR-based automatic decipherment results. FROD improves Top-1 accuracy over OBSD by 3.8 absolute percentage points (42.8% vs. 39.0%). This result indicates that the characters generated by FROD are more likely to be recognized as the correct modern character, suggesting that the proposed font stylization refinement and residual denoising mechanisms improve recognizability and complement the image-quality evaluation. Fig. 4 presents a qualitative comparison. Although OBSD captures the overall structure, it frequently suffers from stroke omissions and misplacements because

FROD

11

Table 1. Comparative evaluation of image translation quality. Method

FID↓

RMSE↓

SSIM↑

LPIPS↓

Pix2Pix DRIT++ CycleGAN BBDM OBSD FROD

246.08 201.77 212.46 122.52 45.18 35.92

0.5302 0.4877 0.4618 0.4181 0.3037 0.2788

0.287 0.313 0.321 0.458 0.622 0.689

0.6037 0.5312 0.5133 0.3978 0.2226 0.1798

Table 2. OCR-based automatic decipherment evaluation of generated images on the character-disjoint augmented dataset. Method

Top-1

Top-5

Top-10

Top-100

Top-500

Pix2Pix DRIT++ CycleGAN BBDM OBSD FROD

0.0% 0.0% 0.0% 18.2% 39.0% 42.8%

0.0% 0.0% 0.0% 22.0% 42.1% 44.0%

0.0% 0.0% 0.0% 23.9% 45.2% 47.2%

1.9% 2.5% 13.8% 33.3% 57.9% 59.1%

6.3% 8.2% 20.1% 37.1% 61.6% 62.9%

of its blind LSS strategy. The comparison between “FROD-w/o Font” and the full FROD output further shows that the font stylization refinement stage suppresses raster artifacts and regularizes stroke appearance. In contrast, FROD, trained with gated segmentation supervision and RDDM, reconstructs intricate stroke topologies with higher fidelity to modern character standards. Fig. 5 also illustrates representative failure cases. The primary limitations arise from incomplete modeling of complex character structures and insufficient recovery of critical local details. In these examples, the preliminary translation preserves only coarse topology, and the stylization stage may amplify structural ambiguities rather than resolve them. These failures explain part of the remaining gap in both image-quality metrics and OCR accuracy. 6.4

Ablation Study Results and Analysis

To isolate and validate the contribution of each component, we conduct ablation studies in Table 3. We consider the following variants: – FROD-AlwaysSeg: Removes the LightGlue threshold and blindly segments all training pairs. – FROD-StandardDiff : Replaces RDDM with a standard noise-predicting conditional diffusion model. – FROD-w/o Font: Removes the multi-stage font stylization refinement step. The results show that “AlwaysSeg” performs substantially worse than the gated strategy, which supports our hypothesis that blindly segmenting highly

12

Y. Hou et al.

Fig. 4. Comparative analysis of oracle bone script image generation quality across different methods.

Fig. 5. Decipherment failure cases.

FROD

13

Table 3. Ablation study of FROD components. Method

FID↓

RMSE↓

SSIM↑

LPIPS↓

FROD-AlwaysSeg FROD-StandardDiff FROD-w/o Font FROD (Full)

44.51 49.33 56.99 35.92

0.3121 0.3418 0.3596 0.2788

0.601 0.582 0.559 0.689

0.2305 0.2511 0.2833 0.1798

abstract OBS pairs damages structural integrity. Replacing RDDM with a standard diffusion model (“StandardDiff”) increases positional drift and degrades FID. Finally, removing the font stylization refinement stage (“w/o Font”) leaves noticeable edge noise and raster artifacts, further confirming that each module is important for producing high-quality translation results. 6.5

Limitations

Although the character-disjoint split evaluates generalization to unseen modern character classes while preventing shared source instances across training and testing, the reported OCR accuracy should still be interpreted as an automatic proxy for isolated-glyph decipherment rather than definitive philological decipherment. End-to-end decipherment ultimately requires contextual, philological, and archaeological evidence beyond isolated glyph images. In addition, the current evaluation reports aggregate performance over the full character-disjoint augmented test set; more fine-grained analysis by stroke complexity, radical composition, and structural deformation will be valuable in future dataset releases.

7

Conclusion

In this paper, we introduced FROD, an image translation framework for assisting the decipherment of oracle bone script by generating recognizable modern Chinese character candidates. To address the severe structural variation between ancient and modern scripts, FROD uses a LightGlue-based feature matching mechanism to provide gated segmentation supervision, thereby improving local radical alignment without forcing mismatched patch correspondences. At inference time, FROD adopts a fixed patch-wise reverse sampling strategy reconstructed with Gaussian blending. We further adapt RDDM to model residual signals explicitly, which mitigates the positional drift and stroke disorder prevalent in standard diffusion methods. Combined with a multi-stage font stylization refinement network, FROD produces clean and standardized modern character images. Extensive experiments show that the proposed method improves image generation quality and downstream OCR accuracy, yielding a 3.8% absolute gain in Top-1 accuracy over OBSD on our character-disjoint augmented dataset. In future work, we plan to explore vector-based font generation for synthesizing scalable outlines and to investigate the integration of multi-task perceptual

14

Y. Hou et al.

recognition losses into diffusion training in order to further improve Top-1 accuracy and domain applicability.

References 1. Guan, H., Yang, H., Wang, X., Han, S., Liu, Y., Jin, L., et al.: Deciphering Oracle Bone Language with Diffusion Models. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15554–15567 (2024) 2. Lindenberger, P., Sarlin, P.-E., Pollefeys, M.: LightGlue: Local feature matching at light speed. In: Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 17627–17638 (2023) 3. Liu, J., Wang, Q., Fan, H., Wang, Y., Tang, Y., Qu, L.: Residual denoising diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2773–2783 (2024) 4. Fu, B., Yu, F., Liu, A., Wang, Z., Wen, J., He, J., et al.: Generate like experts: Multi-stage font generation by incorporating font transfer process into diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 6892–6901 (2024) 5. Guo, J., Wang, C., Roman-Rangel, E., Chao, H., Rui, Y.: Building hierarchical representations for oracle character and sketch recognition. IEEE Transactions on Image Processing 25(1), 104–118 (2016) 6. Huang, S., Wang, H., Liu, Y., Shi, X., Jin, L.: OBC306: A large-scale oracle bone character recognition dataset. In: Proceedings of the International Conference on Document Analysis and Recognition, pp. 681–688 (2019) 7. Li, B., Dai, Q., Gao, F., Zhu, W., Li, Q., Liu, Y.: HWOBC-A handwriting oracle bone character recognition database. Journal of Physics: Conference Series 1651(1), 012050 (2020) 8. Guan, H., Wan, J., Liu, Y., Wang, P., Zhang, K., Kuang, Z., et al.: An open dataset for the evolution of oracle bone characters: EVOBC. arXiv preprint arXiv:2401.12467 (2024) 9. Wang, P., Zhang, K., Wang, X., Han, S., Liu, Y., Wan, J., et al.: An open dataset for oracle bone character recognition and decipherment. Scientific Data 11, 976 (2024) 10. Zhou, J., Tu, Q., Xu, G.: Oracle character recognition using universal inverted bottleneck and inverse image frequency. International Journal on Document Analysis and Recognition 28(4), 609–622 (2025) 11. Isola, P., Zhu, J.-Y., Zhou, T., Efros, A.A.: Image-to-image translation with conditional adversarial networks. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 1125–1134 (2017) 12. Zhu, J.-Y., Park, T., Isola, P., Efros, A.A.: Unpaired image-to-image translation using cycle-consistent adversarial networks. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2223–2232 (2017) 13. Lee, H.-Y., Tseng, H.-Y., Mao, Q., Huang, J.-B., Lu, Y.-D., Singh, M.K., et al.: DRIT++: Diverse image-to-image translation via disentangled representations. International Journal of Computer Vision 128(10–11), 2402–2417 (2020) 14. Saharia, C., Chan, W., Chang, H., Lee, C., Ho, J., Salimans, T., et al.: Palette: Image-to-image diffusion models. In: ACM SIGGRAPH 2022 Conference Proceedings, pp. 1–10 (2022)

FROD

15

15. Li, B., Xue, K., Liu, B., Lai, Y.-K.: BBDM: Image-to-image translation with Brownian bridge diffusion models. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1952–1961 (2023) 16. Zhang, G., Liu, D., Smyth, B., Dong, R.: Deciphering ancient Chinese oracle bone inscriptions using case-based reasoning. In: Sánchez-Ruiz, A.A., Floyd, M.W. (eds.) ICCBR 2021. LNCS, vol. 12877, pp. 309–324. Springer, Cham (2021) 17. Chang, X., Chao, F., Shang, C., Shen, Q.: Sundial-GAN: A cascade generative adversarial networks framework for deciphering oracle bone inscriptions. In: Proceedings of the 30th ACM International Conference on Multimedia, pp. 1195–1203 (2022) 18. Liu, G.-H., Vahdat, A., Huang, D.-A., Theodorou, E.A., Nie, W., Anandkumar, A.: I2 SB: Image-to-image Schrödinger bridge. In: Proceedings of the 40th International Conference on Machine Learning, PMLR 202, pp. 22042–22062 (2023)

Record · ID 919443 · SHA-256 e63ebdc73f90ca4d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.