ConceptioArchivearXiv CS
arXiv CSopen access

cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

cGAP: Generalized Association Plots with HOMALS-Guided Heatmaps for Visualization of High-Dimensional Categorical Data Chun-houh Chen∗1,2 , Shun-Chuan Chang3 , Chiun-How Kao4 , Yi-Ju Lee1 , Shang-Ying Shiu2 , Yin-Jing Tien5 , ShengLi Tzeng6 , Han-Ming Wu7 1 Institute of Statistical Science, Academia Sinica, Taipei, Taiwan 2 Department of Statistics, National Taipei University, New Taipei City, Taiwan

arXiv:2607.15018v1 [stat.ML] 16 Jul 2026

3 Holistic Education Center, Mackay Medical University, New Taipei City, Taiwan 4 Department of Statistics and Data Science, Tamkang University, New Taipei City, Taiwan 5 Institute for Information Industry, Taipei, Taiwan 6 Department of Applied Mathematical and Graduate Institute of Statistics, National Chung

Hsing University, Taiwan 7 Department of Statistics, National Chengchi University, Taipei, Taiwan

Abstract High-dimensional categorical data arise in genetics, biomedicine, and the social sciences, yet visualization tools for such data remain far less developed than those for continuous variables. Existing methods either scale poorly, rely heavily on lowdimensional displays detached from the original data matrix, or prioritize predictive accuracy over interpretability. To address this gap, we introduce categorical Generalized Association Plots (cGAP), a visualization framework for nominal, ordinal, and binary data that preserves the original data matrix while augmenting it with interpretable geometric structure. cGAP uses Homogeneity Analysis (HOMALS) to embed subjects and category levels in a three-dimensional Euclidean space and maps the embedding to red-green-blue coordinates so that similar patterns receive similar colors. The framework integrates three coordinated views: a HOMALSguided heatmap of the raw data matrix, a subject proximity matrix, and a variable proximity matrix. Seriation algorithms are then used to reorder rows and columns to reveal coherent clusters, outliers, and local-to-global structure. We also derive barycentric traceability, projection-distortion, and contrast-preservation properties that clarify how embedding geometry is transferred to the display. We demonstrate the versatility of cGAP through applications to student-animal classification data, mammalian dentition profiles, mushroom records from the UCI Machine Learning Repository, and the Clusters of Orthologous Genes database. These examples show that cGAP supports transparent exploratory analysis by maintaining traceability between derived visual structure and the original categorical observations. cGAP provides a full-matrix, heatmap-based visualization environment for investigating complex categorical datasets across scientific domains. ∗

Corresponding author: [email protected]

1

Keywords: matrix visualization; exploratory data analysis; multivariate categorical analysis; visual analytics; seriation methods

1

Introduction

High-dimensional categorical data arise in many domains, including biology, medicine, text analysis, and the social sciences. In these settings, researchers often seek visual representations that reveal clusters, outliers, and association structure while preserving a clear connection to the original observations. However, visualization tools for multivariate categorical data remain less mature than those for continuous data, especially when the number of variables or categories is large. Existing approaches address this problem from different directions. Classical displays such as bar charts, pie charts, bubble plots, and mosaic plots are useful for low-dimensional categorical data (Hartigan and Kleiner, 1981; Friendly, 1994, 1999), but they become difficult to interpret as dimensionality increases. Embedding-based methods, including homogeneity analysis (HOMALS), multiple correspondence analysis (MCA), and dual scaling, provide low-dimensional geometric representations of categorical structure (Gifi, 1990; Michailidis and de Leeuw, 1998; Greenacre, 1984; Benzecri, 1973; Nishisato, 1984). These optimal scaling concepts have also been adapted for axis-based displays; for example, the textile plot (Kumasaka and Shibata, 2008) extends parallel coordinates by jointly optimizing axis scaling and ordering to align subject trajectories, yielding coordinate representations closely tied to homogeneity analysis. While these methods are valuable for exploratory analysis, the resulting low-dimensional or line-based displays do not by themselves preserve the original data matrix as a clean, scalable primary visual object. In parallel, matrix-based visualization methods such as data images and generalized association plots (GAP) preserve observed values and can reveal large-scale structure effectively (Wegman, 1990; Minnotte and West, 1998; Chen, 2002, 2004; Tien et al., 2008; Wu et al., 2010). Yet categorical data pose an additional challenge: unlike continuous data, their values do not naturally share a common quantitative color scale across variables. This paper introduces categorical Generalized Association Plots (cGAP), a visualization framework for nominal, ordinal, and binary data that combines full-matrix display with interpretable low-dimensional embedding. cGAP uses HOMALS to map subjects and category levels into a three-dimensional Euclidean space, then uses the resulting coordinates to construct color encodings and proximity views for the original matrix. The framework displays three coordinated views: a HOMALS-guided heatmap of the raw data matrix, a subject proximity matrix, and a variable proximity matrix. Through seriation, these views reveal local and global structure while maintaining traceability to the original categorical observations. Our main contributions are as follows. First, we propose a full-matrix visualization framework for high-dimensional categorical data that preserves the original observation table as a central analytic view. Second, we introduce an interpretable color-construction strategy based on HOMALS embeddings, allowing similar categorical patterns to be represented by similar colors. Third, we integrate a raw-data heatmap with subjectproximity and variable-proximity views in a coordinated display that supports the discovery of clusters, outliers, and multilevel association structure. Fourth, we derive traceability, projection-distortion, and contrast-preservation properties that clarify how embedding geometry is transferred to the display. Fifth, we demonstrate the applicability of cGAP 2

across nominal, ordinal, and binary datasets from multiple scientific domains. Figure 1 summarizes graphical tools for continuous data and their categorical counterparts across increasing dimensionality. Within this landscape, cGAP is intended as a high-dimensional categorical analogue of matrix-based data-image visualization, while also incorporating embedding-based information to improve interpretability. The remainder of the paper is organized as follows. Section 2 introduces HOMALS, the measure used to assess the information retained by low-dimensional embedding, and several theoretical properties of the embedding. Section 3 presents the cGAP methodology, including color construction, subject and variable proximity matrices, and contrast enhancement. Section 4 illustrates the method through several applications. Section 5 concludes with discussion and future directions. Dim.

Continuous (Fisher / Iris)

Categorical (Hartigan / Tools) Pie Chart

Bar Chart

Histogram

Holds 13.2%c

Count

No. of Tools

Crank s 4.4% Grips 41.2%

Line Plot

Stacked Bar

Tools

No. of Tools

Petal Width (cm)

Property

Head

Mosaic Plot

Head

Petal Width (cm)

Scatter Plot

Head I

Petal Length Index (cm )

Dim. = 3

Dim. > 3

Dim. = 3

Rotation Plot

Matrix Plot

2D Pie Chart

Dim. > 3 Contingency Plot Matrix (Bubble Plot Matrix)

Bottom

Petal Width (cm)

≥3

Line Plot

(Tool Groups)

Inde x

2

Cuts 23.5%

Property

Petal Width (cm)

1

Turns 10.3% Pounds 7.4%

Head I

Raw Data Matrix Plot

Principle Component Plot

Lengt h

Level 1

Level 2

Level 3

Level 4

Level 5

Optimal Scaling (Homogeneity Analysis)

Grand Tour

PC2 (5.3%)

Dimension Reduction Technique

Index

PL PW

SL

PC1 (92.5%)

SW

GAP (Generalized Association Plot)

?

(Categorical GAP)

Figure 1: Graphical tools for continuous data and their categorical counterparts across increasing dimensionality.

3

2

HOMALS Embedding of Categorical Data

cGAP uses homogeneity analysis (HOMALS) to construct a low-dimensional embedding of categorical observations that supports color encoding and proximity-based views. This section summarizes the components of HOMALS that are needed for the cGAP visualization pipeline.

2.1

Homogeneity Analysis (HOMALS)

Consider J categorical variables observed on N subjects. Let variable Zj have cj categories, and let Gj denote the corresponding N × cj indicator matrix, where Gj (i, t) = 1 if subject i belongs to category t for variable j, and 0 otherwise. HOMALS seeks a joint p-dimensional Euclidean representation for subjects and category levels. We denote the subject coordinates by X ∈ RN ×p and the category coordinates for variable j by Yj ∈ Rcj ×p . The goal is to place subjects near the categories they select and categories near the subjects that select them. This is achieved by minimizing the following average reconstruction loss, where tr(·) denotes the matrix trace: σ(X; Y1 , . . . , YJ ) = J −1

J X





tr (X − Gj Yj )t (X − Gj Yj ) .

(1)

j=1

Without additional constraints, the trivial solution X = 0 and Yj = 0 would minimize the objective. We therefore impose centering and normalization constraints so that the embedding is nondegenerate and the dimensions are comparable, using the N -vector of ones uN and the p × p identity matrix Ip : utN X = 0,

and X t X = N Ip .

(2)

These constraints center the subject configuration at the origin and normalize the embedding. The optimization can be solved by alternating least squares (de Leeuw et al., 1967). Given X, the category coordinates are updated by least squares: Ŷj = (Gtj Gj )−1 Gtj X.

(3)

Given the category coordinates, the subject coordinates are updated as the average of the quantified categories associated with each subject: X̂ = J

−1

J X

Gj Ŷj .

(4)

j=1

X is then re-centered and re-scaled until convergence. The resulting embedding has two properties that are especially useful for cGAP. First, category points lie at the centroids of the subjects assigned to them. Second, subjects with identical response patterns receive identical coordinates. These properties make HOMALS a natural bridge between the original categorical matrix and the Euclidean structure later used for color construction and proximity visualization.

4

2.2

Information Retention in the Embedding

Because cGAP maps the HOMALS solution to three color channels, the quality of a low-dimensional embedding matters directly for visualization. The retained structure can be summarized through the optimized loss, which can be written as:

σ(X; Y1 , . . . , YJ ) = J

−1

J X





tr (X − Gj Yj ) (X − Gj Yj ) = N p − t

p J X X

2 ηjs

(5)

j=1 s=1

j=1 2 with ηjs defined as, where Yj,·s denotes the sth column of Yj : 2 t ≡ Yj,·s Gtj Gj Yj,·s /N. ηjs

(6)

2 Here, ηjs measures how strongly dimension s discriminates among the categories of variable j. The average discrimination power of dimension s across variables is defined as:

γs =

J X

2 ηjs /J.

(7)

j=1

The associated eigenvalues λs = Jγs rank the embedding dimensions by the amount of categorical structure they capture. When p is chosen smaller than the full embedding rank ν of the HOMALS solution, HOMALS acts as a dimension-reduction step. We therefore evaluate the proportion of retained variation using: Pp γs ∗ γp = Pνs=1 . s=1 γs

(8)

This quantity gives the proportion of attainable variation retained by the p-dimensional embedding. In cGAP, it provides a practical check on whether a three-dimensional embedding is adequate for downstream color encoding and structural visualization.

2.3

Theoretical Properties of the Embedding

The HOMALS solution also yields several structural properties that are directly relevant to the interpretability of cGAP. Let cj (i) denote the category selected by subject i for variable j, let x̂i denote the ith row of X̂, and let ŷj,t denote the tth row of Ŷj , that is, the fitted coordinate vector of category t for variable j. Proposition 1 (Barycentric Traceability) For each subject i, the HOMALS coordinate is the barycenter of the category coordinates selected by that subject: x̂i =

J 1X ŷj,c (i) . J j=1 j

(9)

The result follows immediately by taking the ith row of the update equation X̂ = P J −1 Jj=1 Gj Ŷj . Since each row of Gj contains exactly one nonzero entry, the ith row of Gj Ŷj is the coordinate of the category selected by subject i for variable j. This proposition formalizes the traceability of cGAP: each subject point is an average of the categories that define its response profile. Because the RGB relocation used in Section 3.1 is affine, the same identity also holds for displayed colors. 5

Corollary 1 (Affine Color Traceability) For a three-dimensional display, let m denote the maximum absolute coordinate over all fitted subject and category points, and let 13 denote the three-vector of ones. Define the RGB relocation map T (z) = z/(2m) + 0.513 coordinatewise. Then T (x̂i ) =

J 1X T (ŷj,cj (i) ). J j=1

(10)

Hence the displayed color of a subject is the affine average of the colors of its selected categories. This identity justifies the profile-color interpretation used throughout cGAP. Proposition 2 (Projection-Distortion Decomposition) Let ν denote the full (ν) embedding rank of the HOMALS solution. Let x̂i ∈ Rν denote the full-rank HOMALS (p) coordinate of subject i, and let x̂i denote its first p coordinates. Then for any two subjects i and i′ , (ν)

(ν) 2

x̂i − x̂i′

(p)

(p) 2

= x̂i − x̂i′

+

ν X

(x̂is − x̂i′ s )2 .

(11)

s=p+1

Expand the squared Euclidean norm componentwise, where x̂is denotes the sth component (ν) of x̂i , and separate the first p coordinates from the discarded coordinates. Therefore, the p-dimensional display can only underestimate full-space pairwise distances, and the discrepancy is determined exactly by the discarded coordinates. The same decomposition applies to category coordinates. This result complements the retainedvariation criterion γp∗ by giving a geometric interpretation of the distortion induced by restricting the display to three dimensions.

3

Constructing cGAP Displays

This section describes how cGAP transforms a categorical data table into three coordinated visual representations: a HOMALS-guided heatmap of the raw data matrix, a subject proximity matrix, and a variable proximity matrix. We use the animal-grouping data of Nishisato (2007) as a running example. In this dataset, fifteen students grouped thirty-five animals according to perceived similarity, providing a compact illustration of nominal categorical structure.

3.1

Color Encoding from the HOMALS Embedding

The main challenge in extending matrix-based visualization from continuous to categorical data is color construction. For continuous variables, a common numerical scale permits a single ordered color spectrum. For multivariate categorical data, category labels are variable-specific and do not share a global ordering. cGAP addresses this by deriving colors from the geometry of the HOMALS embedding rather than from the labels themselves. This choice is designed to satisfy the relativity principle of statistical graphics (Chen, 2002): subjects and categories that are close in the embedding should receive similar colors, whereas distant configurations should receive more distinct colors. We use a three-dimensional HOMALS embedding because it can be mapped directly to the three channels of a standard display device. RGB is adopted here as a display-native 6

Table 1: Animal-grouping data from fifteen students. Each entry records the group assigned by a student to an animal. The numbers of groups used by the fifteen students were 8, 3, 9, 9 (with no group 9 for Student 4), 7 (Student 5 used groups 0−6), 5, 7, 5, 5, 5, 6, 4, 8, 7, and 7 (with no group 7 for Student 15), producing ninety-five student-group categories in total. Students

Animals

1

2

Dog Alligator Chimpanzee Cow Crow Pigeon Cheetah Chicken Bear Cat Rabbit Frog Goat Tiger Rhinoceros Giraffe Duck Sparrow Hippopotamus Monkey Turkey Pig Crane Leopard Ostrich Lizard Horse Raccoon Tortoise Snake Lion Elephant Camel Hawk Fox

1 1 2 2 3 1 4 1 5 3 5 3 6 1 5 3 1 1 1 1 1 1 2 2 6 1 3 1 7 1 4 1 5 3 5 3 7 1 1 1 5 3 4 1 5 3 4 1 6 3 8 2 1 1 1 1 2 2 2 2 3 1 7 1 7 1 5 3 1 1

3

4

5

6

7

8

9

10 11

12

13 14

15

1 1 2 2 3 3 4 4 5 5 5 5 1 6 5 10 6 4 1 6 7 4 8 9 4 4 1 6 3 4 9 7 5 10 5 5 2 4 3 3 5 10 4 4 5 5 1 6 5 10 8 9 4 4 1 1 8 2 8 2 1 6 2 7 4 7 5 5 1 6

1 2 3 4 5 5 1 5 2 1 1 6 4 1 2 4 5 5 2 3 5 4 5 1 5 6 4 0 6 2 2 4 4 5 1

1 2 1 2 4 4 2 4 2 2 5 3 5 2 5 5 4 4 5 1 4 2 4 2 4 3 1 2 3 3 2 5 5 4 2

1 1 3 3 4 3 3 4 2 5 2 5 3 4 6 5 3 3 1 1 7 1 5 2 7 4 3 4 3 4 3 4 2 5 2 5 3 2 4 3 6 5 7 4 2 5 3 4 3 5 5 2 3 4 7 1 5 2 5 2 3 4 3 4 3 4 2 5 7 1

1 2 3 1 4 4 5 4 5 1 1 2 1 5 5 5 4 4 5 3 4 1 4 5 4 2 1 1 2 2 5 5 5 4 1

1 2 3 2 4 4 1 4 5 1 1 2 1 1 2 3 4 3 2 3 4 1 4 1 4 2 1 1 5 2 1 5 3 4 1

1 2 2 1 4 1 1 4 1 1 1 3 1 1 1 1 4 4 1 2 4 1 4 1 4 3 1 1 3 3 1 1 1 4 1

1 3 3 1 4 4 2 4 1 2 5 6 5 2 7 7 4 4 7 3 4 5 4 2 4 8 5 1 6 8 7 7 7 4 1

1 8 3 4 5 5 2 5 2 1 1 6 1 2 2 4 5 5 2 3 5 1 5 2 4 6 4 1 6 6 2 2 4 5 1

1 2 3 1 4 4 3 1 5 1 1 6 1 2 5 5 4 4 5 3 1 1 4 2 5 6 1 2 6 6 2 2 5 2 2

1 2 2 3 6 6 4 6 4 1 1 7 1 4 5 5 6 6 5 2 6 3 6 4 5 7 3 1 7 7 4 5 5 6 1

coordinate system, not as a claim that it is perceptually optimal among all color spaces. Its practical advantage is that each embedding dimension can be assigned to one channel, allowing the geometry of the embedding to be transferred transparently to color. The analytic content of cGAP therefore lies in relative color similarity and contrast, rather than in any fixed semantic meaning of red, green, or blue individually. As with other low-dimensional embeddings, HOMALS axes are unique up to sign changes and axis permutation. In cGAP, this affects which absolute hues appear in a figure, but it does not alter the underlying proximity structure. Pairwise distances, cluster relationships, and block patterns in the sorted matrices remain unchanged. Throughout this paper, we order the embedding axes by decreasing retained variation (H1, H2, H3) and map them to the R, G, and B channels, respectively. Under this convention, color should be interpreted relationally: similar colors indicate similar embedded configurations, even though a particular hue does not by itself carry a fixed substantive meaning. To map the embedding into display colors, we first relocate subjects and categories into the unit cube [0, 1]3 . Let x ∈ [0, 1]N ×3 and yj ∈ [0, 1]cj ×3 denote the relocated subject and category coordinate matrices, respectively. For subject i, category c of variable j, and 7

RGB channel r ∈ {1, 2, 3}, define x(i, r) =

X̂(i, r) + 0.5 2m

(12)

Ŷj (c, r) + 0.5, (13) 2m where m is the maximum absolute coordinate over all fitted subject and category points before relocation. The first, second, and third coordinates then define the red, green, and blue intensities. This linear rescaling preserves the relative arrangement of points while making them displayable as colors. Points near the center of the configuration are mapped to grayish tones, whereas more extreme configurations move toward the cube faces and corners. Edges connecting subjects to categories inherit the category color, visually linking the joint embedding to the original matrix entries. Figure 2(a) shows the fitted HOMALS embedding, and Figure 2(b) shows the same configuration after relocation into the RGB unit cube. The resulting category colors are displayed in Figure 3(a) and are then used to construct the raw-data heatmap in Figure 3(b). Representative colors for animals are obtained by averaging the colors of the student-group categories associated with each animal (Figure 3(c)), consistent with the dual representation property of HOMALS before optional contrast enhancement. Because cGAP is a color-centered visualization, interpretation should not rely on hue alone. The method is intended to be read jointly through the sorted raw-data heatmap, the subject proximity matrix, the variable proximity matrix, and the associated dendrograms. A more comprehensive study of perceptually optimized palettes and color-vision-deficiency-safe encodings remains an important direction for future work. yj (c, r) =

3.2

Subject Proximity

For subjects, cGAP uses Euclidean distance in the three-dimensional HOMALS embedding as the proximity measure. Figure 3(d) shows the resulting between-animal proximity matrix for the data in Table 1.

3.3

Variable Proximity

For categorical variables, cGAP defines an embedding-based dissimilarity in terms of the HOMALS category coordinates. Let Zk be a categorical variable with ck categories, let ωi denote subject i, and let Zk (ωi ) denote the response of subject i on variable Zk . Define nkt = #{ωi : Zk (ωi ) = t} as the count for category t of variable Zk , and nkl ) = t & Zl (ωi ) = s} as the joint count for categories t and s of variables ts = #{ωi : Zk (ωiP P k Pcl kl k Zk and Zl . Since ct=1 nkt = ct=1 s=1 nts = N , we define the dissimilarity between Zk and Zl as: d(Zk , Zl ) =

N X

D(Zk (ωi ), Zl (ωi )) =

ck X cl X

nkl ts D(t, s),

(14)

t=1 s=1

i=1

where D(t, s) denotes the Euclidean distance between category t of Zk and category s of Zl in the p-dimensional HOMALS embedding. This function d satisfies the four metric conditions: 8

H3

H1 H2

(a) Animal Student_category

(b)

Figure 2: (a) The three-dimensional HOMALS embedding, where H1, H2, and H3 denote the first three dimensions. (b) The same configuration after linear relocation into the unit RGB cube. Edges connect subjects to the categories they select. • Non-negativity: Since D(t, s) ≥ 0, it follows that d(Zk , Zl ) ≥ 0. • Identity of Indiscernibles: If Zk = Zl (i.e., there exists a function g such that g(Zk (ωi )) = Zl (ωi ), ∀ωi ), then the HOMALS solutions for Zk and Zl are identical and hence d(Zk , Zl ) = 0; conversely, if Zk ̸= Zl there exists at least one ωi such that D(Zk (ωi ), Zl (ωi )) > 0, so d(Zk , Zl ) > 0. • Symmetry: The Euclidean distance is symmetric, so d(Zk , Zl ) = d(Zl , Zk ). • Triangle Inequality: Since D(Zk (ωi ), Zl (ωi )) ≤ D(Zk (ωi ), Zm (ωi ))+D(Zm (ωi ), Zl (ωi )) for any third variable Zm and all ωi , it follows that d(Zk , Zl ) ≤ d(Zk , Zm )+d(Zm , Zl ). We interpret d(Zk , Zl ) as a weighted embedding-based dissimilarity, rather than as a strict metric on the original variable space. It is nonnegative and symmetric, and because it aggregates Euclidean distances across subjects it also inherits the triangle inequality. 9

A zero value indicates that the two variables are indistinguishable under this induced comparison for the observed data; it should not be interpreted as semantic identity of the original variables. This level of structure is sufficient for cGAP, whose goal is to organize variables by empirical proximity in the embedding for visualization and seriation. Figure 3(e) displays the resulting between-student proximity matrix for the data in Table 1. S1 S2 S3 S4 S5 S6 S7 S8 S9 S10 S11 S12 S13 S14 S15

(e)

Subject / Variable distance small

large

(a) Dog Alligator Chimpanzee Cow Crow Pigeon Cheetah Chiken Bear Cat Rabbit Frog Goat Tiger Rhinoceros Giraffe Duck Sparrow Hippopotamus Monkey Turkey Pig Crane Leopard Ostrich Lizard Horse Racoon Tortoise Snake Lion Elephant Camel Hawk Fox

(c)

(b)

(d)

Figure 3: The unsorted cGAP display for the animal-grouping data in Table 1 (35 animal samples × 15 student variables): (a) category color map, (b) the raw-data heatmap, (c) animal profile colors, (d) the between-animal Euclidean distance map, and (e) the between-student weighted distance map.

3.4

Integrated Display and Seriation

Having defined the color map and the two proximity matrices, cGAP assembles them into a coordinated display. Figure 3 presents the animal-student data in its original order. Even before reordering, the combined use of the raw-data heatmap and proximity views reveals several recognizable features, including broad animal groupings, heterogeneous treatment of the alligator, the similarity of chimpanzee and monkey, and unusual response patterns for Students S10 and S11. The next step is seriation. GAP algorithms reorder the rows and columns of the rawdata heatmap and the two proximity matrices so that similar subjects and variables appear adjacent. In this paper, we use seriation methods built around Rank-Two Ellipse (R2E), Rank-One Tree (R1T), and the hybrid HCT-R2E procedure. R2E emphasizes smooth global trends, hierarchical clustering tree (HCT) seriation stabilizes local neighborhood structure, and HCT-R2E combines these behaviors by flipping internal HCT nodes under

10

guidance from the R2E ordering. We use average linkage for agglomerative HCT and also consider the analogous R1T-R2E construction for divisive ordering. S10 S11 S2 S1 S6 S12 S14 S13 S7 S15 S9 S4 S3 S5 S8

(e)

(g) Subject / Variable distance small

large

(a) Alligator Hippopotamus Bear Rhinoceros Elephant Lion Cheetah Tiger Horse Camel Giraffe Cow Leopard Dog Racoon Fox Cat Rabbit Pig Goat Ostrich Turkey Chiken Pigeon Hawk Duck Crow Crane Sparrow Lizard Frog Tortoise Snake Monkey Chimpanzee

(c)

A

B

C

D

E

(b)

(d)

(f)

Figure 4: cGAP with HCT-R2E seriation for the animal-grouping data in Table 1 (35 animal samples × 15 student variables): (a) category color map, (b) the sorted raw-data heatmap, (c) animal profile colors, (d) the sorted Euclidean distance map for animals, (e) the sorted weighted distance map for students, (f) the HCT-R2E tree for (d), and (g) the HCT-R2E tree for (e). Figure 4 shows the cGAP results after applying the HCT-R2E seriation algorithm for both the animals (samples) and students (variables). The color patterns in the sorted raw-data heatmap (Figure 4(b)) and animal profiles (Figure 4(c)) identify four main perceptual groups: B. non-primate mammals (sea green), C. birds (pink), D. reptiles and frogs (purple), and E. primates (green). The sorted between-animal distance matrix (Figure 4(d)) clearly highlights the block-diagonal patterns of the four animal groups, featuring a stand-alone cluster (A) of alligators at the top-left. Notably, the alligator elicited diverse classifications, with students assigning it variously to all four possible animal groups or a combination thereof. Figure 4(c) renders the color of the alligator as a blend of purple, sea green, and green. This visual averaging effectively illustrates the principle of relativity in statistical graphics. The ostrich also appears distinct, serving as a bridging animal between the two major clusters of mammals and birds in the proximity matrix (Figure 4(b)), reflecting its categorization with mammals by five students (S1, S7, S11, S14, S15). Meanwhile, the monkey and chimpanzee share nearly identical color profiles and high proximity, indicating they were consistently grouped together by almost all students. Furthermore, the distinct red off-diagonal regions between the primate cluster and other groups in Figure 4(d) visually signify a large distance, confirming their dissimilarity from other animals. The branching patterns of the HCT in Figure 4(d) further delineate the four major animal groups along with the unique relationships of the alligator and ostrich to the other animals. Fundamentally, Figure 4(a) presents the color representation of the categories, where 11

categories sharing compositional similarity are assigned comparable hues. For instance, in Student S8’s classification of 35 animals into five groups, the first, third, and fourth categories, comprising mammals, are rendered in varying shades of sea green. In contrast, the second and fifth categories, which contain non-mammals, are depicted in distinct hues such as purple and pink. Similarly, although Student S5 assigned nine birds to the sixth category and Student S8 assigned them to the fifth, both categories yield identical HOMALS solutions and are consequently displayed in the same pink hue. This confirms that the color representation of categories strictly adheres to the relativity principle.

3.5

Contrast Enhancement for Outlying Configurations

Figures 3 and 4 also illustrate how cGAP exposes atypical samples, variables, and categories. In this dataset, the alligator, the ostrich, and the primate cluster occupy distinctive positions, and Students S10 and S11 show response patterns that differ from those of most other students. Additional local anomalies are also visible, such as the treatment of pigeon by Student S12 and the treatment of sparrow and cow by Student S10. These observations highlight the exploratory role of cGAP: unusual configurations can be identified visually and then investigated substantively in context. A technical challenge arises when outlying configurations lie far from the data center in HOMALS space. After relocation into the unit cube, such points can pull the remaining configurations toward the center, reducing color contrast for the majority of the data. To mitigate this, we apply a power transformation to the relocated coordinates. Let x ∈ [0, 1]N ×3 and yj ∈ [0, 1]cj ×3 denote the relocated subject and category coordinate matrices defined above, and let z = [xt y1t · · · yJt ]t be the resulting M × 3 matrix of display points, where M = N + parameter q ≥ 1, define the transformed matrix z (q) by z (q) (k, r) =

j=1 cj . For a contrast

PJ

z(k, r) − 0.5 1q wk + 0.5, wk

for k ∈ {1, 2, . . . , M }, r ∈ {1, 2, 3}, and wk = max1≤r≤3 |2z(k, r) − 1|. This transformation moves points outward from the center and enhances visual contrast. When q = 1, the display is unchanged because z (1) = z; as q → ∞, points are progressively pushed toward the cube surfaces. In practice, this adjustment improves visual separation while preserving the relational logic of the color encoding. Proposition 3 (Ray-Preserving Contrast Transform) Let 13 be the three-vector of ones and define the centered coordinate uk = z(k, :) − 0.513 . Then the transformed point satisfies 1

z (q) (k, :) − 0.513 = wkq

−1 



z(k, :) − 0.513 .

(15) 1/q−1

Apply the transformation coordinatewise and factor out the common multiplier wk , and the result follows. Because the multiplier is positive, each point remains on the same ray from the center of the RGB cube, its sign pattern relative to the center is unchanged, and any two points on a common ray retain their radial order. Thus the transform changes 12

saturation and visual contrast without altering the directional structure inherited from HOMALS. Together, Table 1 and Figures 2−4 show how cGAP combines embedding, color encoding, proximity structure, and seriation to reveal multilevel categorical patterns in a single exploratory environment. The next section presents larger examples spanning ordinal, nominal, and binary data.

4

Applications Across Data Types

We evaluate cGAP on three datasets chosen to test complementary aspects of categorical visualization: an ordinal biological dataset (mammalian dentition), a nominal benchmark with interpretable class structure (mushrooms), and a large-scale binary genomics dataset (COG profiles). Together, these examples assess how cGAP handles different measurement scales, different levels of structural complexity, and different analysis goals, from compact interpretable clustering problems to large exploratory screening tasks (Table 2). Table 2: Summary of the application examples used to evaluate cGAP. Dataset

Scale

Matrix size

Role in evaluation

What cGAP reveals

Mammalian dentition

Ordinal

66 × 8

Biological ordinal structure

Mushroom attributes

Nominal 8,124 × 22

COG profiles

Binary

Anatomical symmetry, functional dentition gradients, and taxonomy-aligned exceptions Locally decisive variables, edibility regimes, and informative missingness Phylogenetic blocks, complementary present/absent modules, and multiscale genomic structure

4.1

Interpretable class-related structure

2,296 × 5,061 Large-scale exploratory screening

Mammalian Dentition: Ordinal Biological Structure

The mammalian dentition dataset, introduced by Hartigan (1975) and later analyzed with homogeneity analysis by Michailidis and de Leeuw (1998), contains 66 mammals described by eight dentition variables. These variables record the numbers of incisors, canines, premolars, and molars on the upper and lower jaws. Because the variables are ordinal, this example tests whether cGAP can preserve ordered categorical structure while still revealing biologically meaningful groupings. Table 3 lists the full data and the taxonomic order of each mammal. We fit HOMALS with order constraints and obtain a three-dimensional embedding that retains approximately 67% of the attainable variation. Figure 5 shows the resulting cGAP display after HCT-R2E seriation. This example is useful because the dentition 13

variables have clear biological meaning and the taxonomic labels provide an external reference for interpretation. The variable proximity view shows that cGAP recovers coherent ordinal structure rather than isolated variable effects. In particular, the top and bottom versions of the same tooth types lie close together, especially for molars and canines. The category colors also reveal an interpretable opposition between these tooth patterns: mammals with fewer molars tend to retain canines, whereas mammals with more molars often do not. This relationship is expressed simultaneously in the category configuration and in the sorted data matrix. The subject-oriented views further reveal clusters that align well with taxonomy. Carnivora mammals form coherent groups, whereas Rodentia- and Lagomorpha-related mammals occupy a separate region characterized by more molars and fewer canines. Artiodactyla mammals are grouped by the joint absence of incisors and canines. The display also isolates biologically distinctive animals such as the armadillo, whose extreme dentition profile appears as an outlying configuration in both color and proximity space. Overall, Figure 5 shows that cGAP can handle ordered categorical variables without collapsing the data to a single low-dimensional scatterplot, while still linking the observed patterns back to the original table. A further strength of this example is that cGAP reveals not only taxonomic clustering but also structural regularities that cut across taxonomy. The close placement of upper and lower tooth variables indicates that the display captures paired anatomical organization rather than treating each measurement independently. At the sample level, the sorted heatmap suggests a biologically interpretable gradient from carnivore-like dentition, characterized by retained canines and reduced molar counts, toward herbivoreand rodent-like dentition, characterized by reduced canines and relatively expanded molar structure. The display also preserves informative exceptions within this gradient. For example, the Common Mole appears closer to a carnivore-like profile than its taxonomic label alone would suggest, whereas the armadillo forms an extreme boundary case that anchors one end of the dentition space. This combination of symmetry, gradient structure, and localized exceptions illustrates how cGAP can expose functional morphological relationships in ordinal biological data while keeping every observed tooth pattern visible in the raw-data heatmap.

4.2

Mushroom Attributes: Nominal Structure and Edibility

We next apply cGAP to the mushroom dataset from the UCI Machine Learning Repository (Frank and Asuncion, 2010), comprising 8,124 samples from the Agaricus and Lepiota families. Each sample is described by 22 nominal physical-attribute variables together with edibility information. This example tests whether cGAP can expose interpretable nominal structure in a large benchmark where class differences are strong but multivariable relationships remain visually complex. We use a three-dimensional HOMALS solution and HCT-R2E seriation (Tien et al., 2008) to order both samples and variables. Figure 6 presents the coordinated cGAP views. Because many local patterns repeat across mushrooms, the dataset is especially useful for distinguishing variables that contribute little separation from those that organize the major edible and poisonous subgroups. Figure 6 shows that only a subset of the 22 mushroom variables drives the major visual structure. The variable proximity view indicates that attributes such as cap surface, gill 14

Table 3: Dentition data for 66 mammals. Variables are top incisors (TI), bottom incisors (BI), top canines (TC), bottom canines (BC), top premolars (TP), bottom premolars (BP), top molars (TM), bottom molars (BM), and taxonomic order: Didelphimorphia (D), Soricomorpha (S), Chiroptera (Ch), Cingulata (Ci), Lagomorpha (L), Rodentia (R), Carnivora (Ca), and Artiodactyla (A). Mammal

TI BI TC

BC TP BP TM BM

Opossum Hairy Tail Mole Common Mole Star Nose Mole Brown Bat Silver Hair Bat Pigmy Bat House Bat Red Bat Hoary Bat Lump Nose Bat Armadillo Pika Snowshoe Rabbit Beaver Marmot Groundhog Prairie Dog Ground Squirrel Chipmunk Gray Squirrel Fox Squirrel Pocket Gopher Kangaroo Rat Pack Rat Field Mouse Muskrat Black Rat House Mouse Porcupine Guinea Pig Coyote Wolf

3 3 3 3 2 2 2 2 1 1 2 0 2 2 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 3

1 1 0 1 1 1 1 1 1 1 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 1

4 3 2 3 3 3 3 3 3 3 3 0 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 3 3

1 1 1 1 1 1 1 1 1 1 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 1

3 4 3 4 3 2 2 1 2 2 2 0 2 3 2 2 2 2 2 2 1 1 1 1 0 0 0 0 0 1 1 4 4

3 4 3 4 3 3 2 2 2 2 3 0 2 2 1 1 1 1 1 1 1 1 1 1 0 0 0 0 0 1 1 4 4

1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 0

1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1

Order D S S S Ch Ch Ch Ch Ch Ch Ch Ci L L R R R R R R R R R R R R R R R R R Ca Ca

Mammal

TI

BI

TC

BC

TP

BP

TM

BM

Order

Fox Bear Civet Cat Raccoon Marten Fisher Weasel Mink Ferret Wolverine Badger Skunk River Otter Sea Otter Jaguar Ocelot Cougar Lynx Fur Seal Sea Lion Walrus Gray Seal Elephant Seal Peccary Elk Deer Moose Reindeer Antelope Bison Mountain Goat Muskox Mountain Sheep

3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 1 3 2 2 0 0 0 0 0 0 0 0 0

3 3 3 3 3 3 3 3 3 3 3 3 3 2 3 3 3 3 2 2 0 2 1 3 4 4 4 4 4 4 4 4 4

1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 0 0 1 0 0 0 0 0

1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 0 0 0 0 0 0 0 0 0

4 4 4 4 4 4 3 3 3 4 3 3 4 3 3 3 3 3 4 4 3 3 4 3 3 3 3 3 3 3 3 3 3

4 4 4 4 4 4 3 3 3 4 3 3 3 3 2 2 2 2 4 4 3 3 4 3 3 3 3 3 3 3 3 3 3

0 0 0 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 1 1 1 1 1

1 1 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 1 1 1 1 1 1 1 1 1

Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca Ca A A A A A A A A A A

spacing, veil type, ring number, veil color, gill attachment, and cap shape contribute little discrimination, which is consistent with their relatively monotone appearance in the data matrix. By contrast, gill color, ring type, stalk color above ring, odor, and habitat show strong contrast and organize the main sample blocks, producing six visually coherent mushroom groups in the coordinated views. These groups align closely with edibility patterns. Some clusters are almost entirely edible or poisonous, whereas mixed groups can be interpreted through a small number of variables, especially odor and habitat. The visualization also makes the missing-data structure transparent: missingness occurs only in the stalk-root variable and is closely associated with ring-type and ring-number patterns, suggesting a systematic rather than random mechanism. Figure 7 summarizes one set of rules derived from this visual analysis. cGAP is not introduced here as a predictive classifier; rather, the mushroom example shows how the display can reduce a 22-variable nominal dataset to a smaller set of visually grounded and substantively interpretable decision cues. This example also highlights a less obvious capability of cGAP: it distinguishes globally weak variables from locally decisive ones. Variables such as veil type or cap shape appear visually flat across most of the matrix, immediately identifying them as poor drivers of subgroup separation, whereas odor, habitat, gill color, and ring-related variables become decisive only within specific blocks of mushrooms. In that sense, the display does more than separate edible from poisonous samples; it reveals multiple local rule regimes for edibility rather than a single global decision boundary. The mixed clusters are especially 15

4 3 2 1 0

Mammals

Taxonomic Order

(e)

Ocelot Cougar Lynx Jaguar Walrus Elephant.Seal Weasel Mink River.Otter Sea.Otter Grey.Seal Skunk Badger Ferrer Civet.Cat Fur.Seal Sea.Lion Wolverine Fisher Marten Field.Mouse Muskrat Pack.Rat House.Mouse Black.Rat Pocket.Gopher Fox.Squirrel Gray.Squirrel Kangaroo.Rat Guinea.Pig Porcupine Chipmunk Beaver Marmot Groundhog Prairie.Dog Ground.Squirrel Pika Snowshoe.Rabbit Armadillo Raccoon Wolf Fox Bear Common.Mole Opposum Coyote Hairy.Tail.Mole Star.Nose.Mole Peccary Brown.Bat Lump.Nose.Bat Silver.Hair.Bat Red.Bat Hoary.Bat Pigmy.Bat House.Bat Reindeer Elk Deer Moose Antelope Bison Mountain.Goat Muskox Mountain.Sheep

(h)(c)

(a)

Order

BM TM BC TC TP BP BI TI

(g)

Didelphimorphia Soricomorpha Chiroptera Cingulata Lagomorpha Rodentia Carnivora Artiodactyla

Subject / Variable distance small

large

A

B

C

D

E

(b)

(d)

(f)

Figure 5: cGAP analysis of the mammalian dentition data (66 mammals × 8 ordinal variables): (a) category color map, (b) sorted raw-data heatmap, (c) mammal profile colors, (d) sorted Euclidean distance map for mammals, (e) sorted weighted distance map for dentition variables, (f) HCT-R2E for (d), (g) HCT-R2E for (e), and (h) mammal names with taxonomic-order color legend. informative because they show that cGAP can expose alternative pathways to similar class labels, such as habitat-based separation in one block and odor-based separation in another. The structured missingness in stalk-root strengthens this interpretation by showing that cGAP can simultaneously function as a pattern-discovery tool and a dataaudit tool, revealing when apparently secondary variables are tied to systematic recording or biological substructure.

4.3

COG Profiles: Large-Scale Binary Structure

Our third example evaluates cGAP on large-scale binary data using the Clusters of Orthologous Genes (COG) database (Tatusov et al., 1997, 2000, 2001; Galperin et al., 2015, 2021, 2024). We analyze 2,296 representative genomes (2,103 bacteria and 193 archaea) and 5,061 COG variables. Although the original database records paralog counts, we convert the table to a binary presence/absence matrix so that each entry indicates whether a genome contains a given COG. These phyletic profiles are widely used to study conserved functions, lineage-specific gene loss, and horizontal transfer (Glazko and Mushegian, 2004; Tzeng et al., 2009). Within cGAP, the two binary states are treated as nominal categories and embedded jointly with genomes and COG variables. Figure 8 emphasizes the matrix-scale view of the data: genomes retain their taxonomic ordering, whereas COGs are reordered by the R1T-R2E seriation algorithm. To keep the main figure readable at this scale, the primary panel focuses on the raw-data heatmap and the screening display; full-resolution proximity views are better suited to supplementary material in a journal presentation. Figure 8 demonstrates how cGAP scales to large binary data while preserving a 16

A B

1 2 3 45 6 7

89

1011

12

A B B

C

D

large E

F

8

9

yellow white orange brown

spicy pungent none musty foul fishy creosote anise almond

odor

stalk_color_above_ring

ring_type

two one none

5

yellow white red pink orange gray cinnamon buff brown

6 stalk_root

7

stalk_color_below_ring

Subject / Variable distance

small

smooth silky scaly fibrous

4

yellow white red pink orange gray cinnamon buff brown

veil_color

(g)

3

pendant none large flaring evanescent

ring_number

(a)

1

2

yellow white purple orange green chocolate buff brown black

stalk_surface_above_ring

edible / poisonous

(e)

yellow white red purple pink orange green gray chocolate buff brown black

spore_print_color

stalk_surface_below_ring bruises cap_surface gill_spacing veil_type 8 ring_number 9 veil_color gill_attachment cap_shape 10 stalk_root 11 habitat stalk_shape gill_size 12 population cap_color

gill_color

1 gill_color 2 spore_print_color 3 ring_type 4 stalk_color_above_ring 5 stalk_color_below_ring 6 odor 7 stalk_surface_above_ring

10

rooted equal club bulbous missing

A B

C

C

C

D

D

D

E

E

E

F

F

F

11

f

(c)

1 2 3 45 6 7

89

(b)

1011

12

B

C

D

E

(d)

woods waste urban paths meadows leaves grasses

population

(f)

habitat

(h)

12

solitary several scattered numerous clustered abundant

F

Figure 6: cGAP analysis of the mushroom data (8,124 mushrooms × 22 nominal variables): (a) category color map, (b) sorted raw-data heatmap, (c) mushroom profile colors, (d) sorted Euclidean distance map for mushrooms, (e) sorted weighted distance map for variables, (f) HCT-R2E for (d), (g) HCT-R2E for (e), and (h) edibility labels for the samples. readable matrix-centered view. ARCHAEA form a distinct block characterized by several archaeal COG groups, revealing both coarse separation from bacteria and finer internal subdivision, including multiple EURYARCHAEOTA clusters. The a and a groups highlight an important advantage of cGAP for binary data: present and absent states can have inverse but closely related structure, so the method assigns them contrasting colors while keeping them close in the embedding. BACILLOTA and PSEUDOMONADOTA also show strong lineage-specific signatures, with the b1 − b4 and d1 − d5 groups differentiating major classes and connecting these blocks to secretion-, motility-, and membrane-related functions. The screening view in Figure 8(b) reveals additional structure that is harder to see in the full matrix, including the association of BACTEROIDOTA with the c1 group, near-universal w1/w3 patterns, and genomes with near-total depletion of annotated COGs. This screening panel behaves like a presence-only overlay, but it is not equivalent to a conventional binary GAP based on Jaccard-type distances, because cGAP keeps complementary inverse patterns analytically connected through the HOMALS embedding. Figure 9 complements the matrix display with sediment plots summarizing marginal presence frequencies for COGs and genomes, showing that the method supports both local structural interpretation and global distributional screening in a large genomic application. The COG application is particularly persuasive because it shows that cGAP can 17

stalk color above ring 3. cinnamon

ring type

gill color

5. orange

others

4. large

others

3. buffed

others

36 / 0

0 / 192

others

1296 / 0

others

1728 / 0

others

Group A

Group B

Group C

Group D

population 2. clustered

others

Group E

Group F

habitat

odor

6. waste

2. leaves 1. almond 2. anise

0 / 192

16 / 0

7. none

0 / 800

others (3, 5, 8)

736 / 0

habitat

7. woods

others

stalk surface above ring

spore print color

3. silky

others

5.green

32 / 0

0 / 1784 72 / 0

others

0 / 1240

Figure 7: One sequence of interpretable rules, derived from the cGAP views, for distinguishing poisonous from edible mushrooms using a reduced set of physical-attribute variables. represent several layers of genomic structure in a single coordinated analysis. At the coarsest level, the raw-data heatmap preserves broad phylogenetic organization, separating archaeal and bacterial regimes. Within those regimes, the reordered COG blocks expose pathway- and function-level modules that refine the taxonomic picture, such as archaeal ribosomal signatures, Bacillota-specific cell-cycle patterns, cyanobacterial photosystemrelated groups, and secretion- or motility-associated patterns within Pseudomonadota. The inverse a and a relationships are especially important methodologically: they demonstrate that cGAP does not reduce binary structure to simple presence counts, but can keep complementary present and absent states close in the embedding while rendering them visually distinct. The screening view and sediment displays then add a multiscale perspective, making near-universal cores, lineage-specific systems, and genomes with unusually sparse annotations visible within the same analytic framework. Together, these features make the COG study a strong demonstration that cGAP remains interpretable even when the data are both high dimensional and biologically heterogeneous. Across these three applications, cGAP shows consistent strengths: it preserves the raw categorical matrix, reveals subject- and variable-level structure through coordinated views, and remains interpretable across ordinal, nominal, and binary settings. More specifically, the applications show that cGAP can recover anatomical symmetry and functional gradients in ordinal data, distinguish globally weak variables from locally decisive ones in nominal data, and expose complementary present/absent modules together with multiscale phylogenetic structure in large binary data. Taken together, these examples suggest a practical presentation strategy for journal submission: the main paper can emphasize the integrated matrix displays and central findings, while large supporting proximity views can be provided as supplementary material when space is limited.

18

CRENARCHAEOTA EURYARCHAEOTA

a1 a2 a2 a3 a4 a5 a6

c1

Ribosome 30S subunit Ribosome 50S subunit Archaeal ribosomal proteins Type VI secretion system Photosystem I Photosystem II Type IX secretion/gliding motility system

c2

b1

d1

functional category

a1

pathway

parent taxonomic pathway functional category

ARCHAEA BACTERIA BACILLOTA PSEUDOMONADOTA

D. Cell cycle control, cell division, chromosome partitioning C. Energy production and conversion J. Translation, ribosomal structure and biogenesis M. Cell wall/membrane/envelope biogenesis G. Carbohydrate transport and metabolism U. Intracellular trafficking, secretion, and vesicular transport L. Replication, recombination and repair N. Cell motility

d2

d3

b2 b3 b4

d4

d5

A

BACILLI CLOSTRIDIA

B

ACTINOMYCETOTA

BACTEROIDOTA

C

CYANOBACTERIOTA MYCOPLASMOTA

OTHER_BACTERIA

ALPHAPROTEOBACTERIA

BETAPROTEOBACTERIA

D

GAMMAPROTEOBACTERIA

parent taxonomic category taxonomic category pathway functional category

(a) w1 w2 w3c1

A B

BACTEROIDOTA

C

MYCOPLASMOTA

OTHER_BACTERIA

ALPHAPROTEOBACTERIA

D GAMMAPROTEOBACTERIA

(b)

Figure 8: cGAP analysis of the COG dataset (2,296 genomes × 5,061 variables): (a) rawdata heatmap with genomes kept in taxonomic order and COGs reordered by R1T-R2E seriation, and (b) a screening view that masks absent states and shows only present COGs. Only selected pathways and functional categories discussed in the text are labeled for readability.

5

Discussion and Conclusion

This paper presents cGAP, a full-matrix visualization framework for high-dimensional categorical data that combines HOMALS embedding, color encoding, proximity matrices, and seriation in a coordinated display. The main contribution of cGAP is not simply to produce a low-dimensional view of categorical structure, but to preserve the original data matrix as a central analytic object while augmenting it with interpretable geometric and relational context. Across the examples in this paper, cGAP reveals clusters, outliers, complementary category patterns, and variable groupings in nominal, ordinal, and binary data. The method is particularly useful when analysts need to move repeatedly between derived structure and raw observations. In contrast to standalone scatterplots or dimensionreduction displays, cGAP allows users to inspect every matrix entry while still benefiting from embedding-based color similarity and proximity-based organization. This makes the 19

A B

C

D

(a)

(b)

(c)

(d)

Q1

M

1706

Q3

405

1244

1508

1769

168

Q1

M

(e)

(f)

Q3

Figure 9: Sediment displays for the COG dataset (2,296 genomes × 5,061 COGs): (a) COG-wise and (b) genome-wise plots following the arrangement in Figure 8(a), (c) and (d) the corresponding plots reordered by presence frequency, and (e) and (f) black-and-white presence/absence views of (c) and (d). framework well suited to exploratory analysis tasks in which traceability, local pattern inspection, and multilevel comparison are all important. The barycentric traceability identity, projection-distortion decomposition, and ray-preserving contrast transform directly clarify why these displays can be interpreted in a principled way. At the same time, cGAP has several limitations. First, the visual quality of the color encoding depends on how much categorical structure is retained by the chosen low-dimensional HOMALS solution. When the retained variation in three dimensions is limited, the color map may underrepresent some relationships even though the matrix display itself remains available. Second, RGB is used here as a practical display-native coordinate system rather than as a perceptually optimal color space. Although this design preserves the geometry-to-color mapping transparently, more perceptually calibrated and color-vision-deficiency-aware encodings would be valuable extensions. Third, very large datasets may require selective labeling, interactive exploration, or supplementary highresolution views to fully expose subject and variable proximity structure in publication figures. These limitations also suggest several directions for future work. One direction is to study alternative color spaces and constrained mappings that better balance interpretability, perceptual uniformity, and accessibility. Another is to extend cGAP with interactive interfaces that support zooming, brushing, linked highlighting, and user-controlled filtering for large categorical matrices. It would also be useful to investigate hybrid workflows in which higher-dimensional HOMALS solutions are used for proximity analysis while three-dimensional projections are reserved for display color construction. 20

In conclusion, cGAP provides an interpretable visualization environment for highdimensional categorical data by linking the original matrix to embedding-based color structure and proximity-based organization. Rather than replacing existing multivariate methods, it complements them by offering a matrix-centered view that preserves data traceability while supporting exploratory discovery. We expect cGAP to be especially useful in domains where categorical structure is complex, high-dimensional, and scientifically important, including biomedical, genomic, and social-science applications.

Acknowledgment The authors thank the late Dr. Chih-Wen Ou-yang, a former postdoctoral fellow and research assistant, for invaluable discussions and insightful contributions to the development of the cGAP methodology. His dedication and intellectual input greatly shaped this work. This work was supported by the Ministry of Science and Technology, Taiwan, under the grant “A comparison of matrix visualization and EDA of non-quantitative data” (MOST 107-2118-M-001-004-MY2).

Software Availability The software described in this study is available for download at https://gap.stat. sinica.edu.tw/software.html and can also be accessed via an online version at https: //maokao.github.io/cGAPOnline/.

References Benzecri, J.P., 1973. L’analyse des donnees. II. L’analyse des correspondances. Dunod, Paris. Chen, C.H., 2002. Generalized association plots: information visualization via iteratively generated correlation matrices. Statistica Sinica 12, 7–29. Chen, C.H., 2004. Matrix visualization and information mining, in: COMPSTAT 2004 – Proceedings in Computational Statistics, pp. 85–100. Frank, A., Asuncion, A., 2010. Uci machine learning repository. URL: http://archive. ics.uci.edu/ml. irvine, CA: University of California, School of Information and Computer Science. Friendly, M., 1994. Mosaic displays for multi-way contingency tables. Journal of the American Statistical Association 89, 190–200. Friendly, M., 1999. Extending mosaic displays: Marginal, conditional, and partial views of categorical data. Journal of Computational and Graphical Statistics 8, 373–395. Galperin, M.Y., Makarova, K.S., Wolf, Y.I., Koonin, E.V., 2015. Expanded microbial genome coverage and improved protein family annotation in the COG database. Nucleic Acids Research 43, D261–D269.

21

Galperin, M.Y., Wolf, Y.I., Makarova, K.S., Vera Alvarez, R., Landsman, D., Koonin, E.V., 2021. COG database update: focus on microbial diversity, model organisms and widespread pathogens. Nucleic Acids Research 49, D274–D281. Galperin, M.Y., et al., 2024. COG database update 2024. Nucleic Acids Research 53, D356– D363. URL: https://doi.org/10.1093/nar/gkae983, doi:10.1093/nar/gkae983. Gifi, A., 1990. Nonlinear Multivariate Analysis. John Wiley & Sons. Glazko, G.V., Mushegian, A.R., 2004. Detection of evolutionarily stable fragments of cellular pathways by hierarchical clustering of phyletic patterns. Genome Biology 5, R32. URL: https://doi.org/10.1186/2004-5-5-r32, doi:10.1186/2004-5-5-r32. Greenacre, M.J., 1984. Theory and Applications of Correspondence Analysis. Academic Press, London. Hartigan, J., 1975. Clustering Algorithms. Wiley, New York, NY. Hartigan, J.A., Kleiner, B., 1981. Mosaics for contingency tables, in: Eddy, W.F. (Ed.), Computer Science and Statistics: Proceedings of the 13th Symposium on the Interface. Springer-Verlag, New York. Kumasaka, N., Shibata, R., 2008. High-dimensional data visualisation: The textile plot. Computational statistics & data analysis 52, 3616–3644. de Leeuw, J., Young, F.W., Takane, Y., 1967. Additive structure in qualitative data: An alternating least squares method with optimal scaling features. Psychometrika 41, 471–503. Michailidis, G., de Leeuw, J., 1998. The Gifi system of descriptive multivariate analysis. Statistical Science 13, 307–336. Minnotte, M., West, R.W., 1998. The Data Image: A tool for exploring high dimensional data sets, in: Proceedings of the Section on Statistical Graphics, American Statistical Association, Alexandria, Virginia. Nishisato, S., 1984. Dual scaling of reciprocal medians, in: Proceedings of the 32nd Scientific Conference of the Italian Statistical Society, Societa Italiana di Statistica, Sorrento, Italy. pp. 141–147. Nishisato, S., 2007. Multidimensional Nonlinear Descriptive Analysis. Chapman & Hall. Tatusov, R.L., Galperin, M.Y., Natale, D.A., Koonin, E.V., 2000. The COG database: a tool for genome-scale analysis of protein functions and evolution. Nucleic Acids Research 28, 33–36. Tatusov, R.L., Koonin, E.V., Lipman, D.J., 1997. A genomic perspective on protein families. Science 278, 631–637. Tatusov, R.L., Natale, D.A., Garkavtsev, I.V., Tatusova, T.A., Shankavaram, U.T., Rao, B.S., Kiryutin, B., Galperin, M.Y., Fedorova, N.D., Koonin, E.V., 2001. The COG database: new developments in phylogenetic classification of proteins from complete genomes. Nucleic Acids Research 29, 22–28. 22

Tien, Y.J., et al., 2008. Methods for simultaneously identifying coherent local clusters with smooth global patterns in gene expression profiles. BMC Bioinformatics 9, 1–16. Article 155. Tzeng, S.L., Wu, H.M., Chen, C.H., 2009. Selection of proximity measures for matrix visualization of binary data, in: Proceedings of the 2009 2nd International Conference on BioMedical Engineering and Informatics (BMEI 2009), pp. 1932–1940. Wegman, E.J., 1990. Hyperdimensional data analysis using parallel coordinates. Journal of the American Statistical Association 85, 664–675. Wu, H.M., Tien, Y.J., Chen, C.H., 2010. GAP: A graphical environment for matrix visualization and cluster analysis. Computational Statistics and Data Analysis 54, 767–778.

23

Record · ID 373389 · SHA-256 74287bde1ed32c97
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.