Upgrade to Pro — share decks privately, control downloads, hide ads and more …

Tensor decomposition-based unsupervised feature...

Sponsored · Your Podcast. Everywhere. Effortlessly. Share. Educate. Inspire. Entertain. You do you. We'll handle the rest. →
Avatar for Y-h. Taguchi Y-h. Taguchi PRO
September 30, 2026

Tensor decomposition-based unsupervised feature extraction applied to bioinformatics and affinity with AI-based methodology

BIC (Bioinformatics Center Seminar, Kyoto Univ.)
1st Oct 2026

Avatar for Y-h. Taguchi

Y-h. Taguchi PRO

September 30, 2026

More Decks by Y-h. Taguchi

Other Decks in Science

Transcript

  1. Tensor decomposition-based unsupervised feature extraction applied to bioinformatics and affinity

    with AI-based methodology Y-h. Taguchi Department of Physics, Chuo University (will stay here till the end of March, 2027. Room 220B) BIC seminar 2026.10.1 1
  2. Brief Self Introduction Graduated from Science Tokyo (formerly, Tokyo Inst.

    Tech.) Bs (Dept. Applied Phys, 1984), Ms & Dr (Dept. Physics, 1986 & 1988) (A part time job at Sunrise, Gundam company) Mainly Bioinformatics, multiomics analysis using Tensor One edition from MIMB. BIC seminar 2026.10.1 2
  3. Many popular science books (mostly in Japanese, some were translated

    into Korean and Chinese) BIC seminar 2026.10.1 3
  4. What is tensor? Matrix Tensor xij xijk Gene i ×

    Subject j × Methylation k Gene i × Subject j Gene i × Subject j × Time k BIC seminar 2026.10.1 5
  5. What is tensor decomposition? Singular value decomposition x ij ∼∑

    λ u M N xij l l li v lj L ~ N × L uli M L ×L vlj λl BIC seminar 2026.10.1 6
  6. Tensor decomposition (TD) x ijk ∼∑ λ u l l

    li v lj Kw K N xijk ~ Nu w lk M 1k v1j 1i Kw + N u M 2k v2j +··· 2i M BIC seminar 2026.10.1 7
  7. x ijk ∼∑ l 1 ,l u 2 l 1

    i v l 1 l 2 j L1 K N xijk ~N ul1i w l 2 k L2 × L1 M M vl1l2j × L2 wl2k M BIC seminar 2026.10.1 8
  8. x ijk ∼∑ K N xijk M l 1 ,l

    2 ,l G 3 ~ (l l l )u 1 2 3 L3 G L 2 L1 N ul1i l 1 i u l 2 j u l 3 k K ul3k ul2j M Tucker Decomposition ← employed! BIC seminar 2026.10.1 9
  9. TD based unsupervised Feature Extraction (FE) x j 1 j

    2 ⋯j L i β Lβ Sample conditions (subject, tissue, time, sex, age, drug, etc) 1 i 2 ⋯i L α Lα Features (gene, methylation, microRNA, etc) Extraction of subsets of features, critical to sample conditions BIC seminar 2026.10.1 10
  10. A typical question: “Which gene (i) expressions are distinct between

    patients and healthy control (j1) in tissue (j2) specific manners with the treatment of which combination of drugs (j3)?” → One feature vs three sample conditions x j 1 j 2 j 3 i ∈ℝ M 1 ×M ×M ×N 2 BIC seminar 2026.10.1 3 11
  11. How to select genes? 1. Investigate ul1j1, ul2j2, ul3j3 2.

    Find ul1j1 distinct between patients and healthy control, tissue and drug specific ul2j2, ul3j3 3. Find l4 associated with the absolutely largest G(l1 l2 l3 l4) 4. Extract a set of i with larger absolute ul4i. BIC seminar 2026.10.1 12
  12. ul1j1 ul2j2 Healthy control Tissues specific j1 j2 patients ul3j3

    l1,l2,l3 selected G(l1 l2 l3 l4) BIC seminar 2026.10.1 l4 Drugs specific j3 13
  13. Selection of i using ul4i We assume …. is to

    be selected P(ul4i) Null hypothesis = Gaussian for not selected i Large Pi P i =P χ 2 [ ( )] > u 2 l σ 4 i Small Pi True σl l Overestimated σl ul4i Smaller selected i Larger, less significant Pi BIC seminar 2026.10.1 14
  14. Optimization of σl Histogram of 1-Pi Flatness of h(1-P): σh

    (excluding is to be selected) Idealized case σh σh BIC seminar 2026.10.1 σh 15
  15. How to select genes? 1. Investigate ul1j1, ul2j2, ul3j3 2.

    Find ul1j1 distinct between patients and healthy control, tissue and drug specific ul2j2, ul3j3 3. Find l4 associated with the absolutely largest G(l1 l2 l3 l4) 4. Extract a set of i with larger absolute ul4i. BIC seminar 2026.10.1 17
  16. Practical example: Subjects (mice) × Tissues × Drugs Taguchi Y

    and Turki T (2020) Universal Nature of Drug Treatment Responses in DrugTissue-Wide Model-Animal Experiments Using Tensor Decomposition-Based Unsupervised Feature Extraction. Front. Genet. 11:695. doi: 10.3389/fgene.2020.00695 BIC seminar 2026.10.1 18
  17. Drugs (1) Alendronate, (2) APAP, (3) Aripiprazole, (4) Asenapine, (5)

    Cisplatin, (6) Clozapine, (7) Dox, (8) EMPA, (9) FivePercentSucrose, (10) Lenalidomide, (11) Lurasidone, (12) Olanzapine, (13) Repatha, (14) Risedronate, (15) Sofosbuvir, (16) Teriparatide, (17) WT.No.treated, (18) 5percentCMC0.25percentTween80. BIC seminar 2026.10.1 20
  18. Optimized σl increase the number of selected is Taguchi, Yh.,

    Turki, T. Adapted tensor decomposition and PCA based unsupervised feature extraction select more biologically reasonable differentially expressed genes than conventional methods. Sci Rep 12, 17438 (2022). https://doi.org/10.1038/s41598-02221474-z BIC seminar 2026.10.1 21
  19. In some sense, it is fully-AIautomated research to check whether

    AI can perform TD-based unsupervised FE instead of me BIC seminar 2026.10.1 25
  20. To GPT-6 sol: に関係するsingle cell ベースのマルチオミックス手法で新しいもの はありますか? YT:“genomic regulation (Are

    there any new single-cell-based multi-omics approaches related to genomic regulation? ) ” Here after only English translation is presented. BIC seminar 2026.10.1 26
  21. YT:“Please create a table showing the samples and modal types

    for GSE303006.” BIC seminar 2026.10.1 28
  22. GPT: “Therefore, if we are to perform a new analysis

    using TD in this study, it would be more innovative to generate a broader set of candidate E–P pairs ourselves using the method described above and let the TD simulation select the E–P pairs, rather than simply using the 7,146 E–P pairs from CHARM as-is. In particular, I find the proposal to create to be the most interesting. This fits quite nicely with the GSE303006 dataset in this study.” BIC seminar 2026.10.1 29
  23. Since replication of the conversation with GPT-6.sol is a bit

    confusing, I will not replicate all the conversation with GPT-6.sol, but I proposed almost nothing although I have sometimes commented about something wrong. Since I have published numerous papers (with open access) often with Github-deposited R-code, GPT-6.sol could learn almost everything there and made suitable decisions. BIC seminar 2026.10.1 30
  24. After quality control, 730,969 E–P pairs, 42,669 enhancers, 16,239 genes,

    391,435 20-kb bin pairs, and 4,258 cells were retained (by GPT-6.sol). BIC seminar 2026.10.1 31
  25. Normalization For ATAC and H3K27ac, cell depth was calculated using

    all enhancer candidate rows, and the retained enhancer matrix was transformed as where xfc is the fragment count for feature f in cell c and Dc is the corresponding total depth. BIC seminar 2026.10.1 32
  26. The RNA matrix was normalized to the same logCP10K scale.

    Each retained feature was then centered and scaled across cells. BIC seminar 2026.10.1 33
  27. For 3D, positive finite Euclidean distances dbc were converted to

    proximity-like values For each 20-kb bin pair, qbc was standardized over observed cells. Missing standardized 3D values were set to zero, so that missingness corresponded to the feature mean after standardization. BIC seminar 2026.10.1 34
  28. Stage 1: Each modality was reduced independently to K =

    20 components. For ATAC, H3K27ac, 3D distance and RNA, truncated SVD was applied to the feature-wise standardized matrix. The decomposition was expressed as [Features (B) vs cells (V)] where Bm contains modality-specific feature scores and Vm contains orthonormal cell loadings. ATAC and H3K27ac feature rows correspond to enhancers, RNA rows to genes, and 3D rows to 20-kb E–P bin pairs (see P29). BIC seminar 2026.10.1 35
  29. Stage 1: produced feature-score matrices, Bm, of 42,669 × 20

    for ATAC, • 42,669 × 20 for H3K27ac, • 16,239 × 20 for RNA, • 391,435 × 20 for 3D. • The corresponding cell loading matrices were numerically orthonormal, with orthogonality errors on the order of 10−14. BIC seminar 2026.10.1 36
  30. The conceptual stage-2 tensor was Its entries were defined as

    follows: B[*(p),k] is the summation of Bm with each e-p pair. d-1/2*(p) is scaling constant. BIC seminar 2026.10.1 37
  31. x pckm ∼∑ l 1 ,l 2 ,l 3 ,l

    G 4 (l l l l )u 1 2 3 4 l 1 p u l 2 c u l 3 k u l 4 m HOSVD was performed using Tucker ranks 20 × 20 × 20 × 4 for the E–P, cell, stage-1component, and modality modes, respectively. The final 20 × 20 × 20 × 4 Tucker core retained 58.41% of the conceptual tensor BIC seminar 2026.10.1 38
  32. Largest 100 G included 12 unique ul2c. Correlations between 12

    ul2c and cell labels or three replicates are computed. They are more correlated with cell labels. G(3, 3, 2, 1) is the best. BIC seminar 2026.10.1 39
  33. Enrichment analysis For every E–P pair p, the signed E–P-side

    score associated with this core element was To produce a signed ranking of all 16,239 genes, a degree-corrected target-gene score was then calculated from all 730,969 E–P pairs as follows: BIC seminar 2026.10.1 40
  34. Of the 16,239 ranked genes, 14,102 (86.84%) mapped to the

    mouse annotation database used for GO analysis. Preranked GO Biological Process GSEA returned 5,495 terms. At FDR ≤ 0.05, 112 terms had positive NES and 28 had negative NES. Reasonable terms (see next pages) BIC seminar 2026.10.1 41
  35. After submissions…… It was invited by Prof. Nakai to be

    included into special issue of Genes (MDPI), but editorial office refused to apply free APC. Then withdrawn. Communication biology desk rejected it and transferred to Sci. Rep. that also desk rejected. BMC Bioinformatics did not desk rejected and sent it to reviews. Let us see what will go on…. BIC seminar 2026.10.1 43
  36. I have never been rejected by Genes (MDPI) for the

    invitation of APC free submission. Sci. Rep. has never desk rejected mine and the reason is funny. “Among the considerations that arise at this stage is the degree to which the results will stimulate new thinking in the field. In this case, we find that the manuscript does not represent a sufficiently valid or original finding in our understanding of Tensor Decomposition of Single-Cell Four-Omics Data Reveals Cell-Type-Associated Enhance, and we are therefore unable to consider it further.” BIC seminar 2026.10.1 44
  37. Conclusions TD based unsupervised FE was proposed and applied to

    many bioinformatics problems. AI can now perform TD based unsupervised FE, but possibly because of that, the research might not be valid enough to be accepted? BIC seminar 2026.10.1 45