Slide 1

Slide 1 text

Retrieval of LoRA Models based on Layer-Wise Weight Embedding without Metadata (Shizuoka University) (University of Tsukuba) (University of Tsukuba, NII) (University of Hyogo) (LY Corporation) (Shizuoka University) 16th ACM International Conference on Multimedia Retrieval Yuma OE Huu-Long PHAM (Shizuoka University) Yuro KANADA Makoto P. KATO Hiroaki OHSHIMA Sumio FUJITA Yoshiyuki SHOJI Brave New Ideas

Slide 2

Slide 2 text

Need to Efficiently Manage a Large Number of LoRA Models! 1 Make it possible to calculate model similarity even without metadata or output samples! Need to manage a large number of LoRA models !! Logo LoRA Brand Logo Design own image-genAI Manual Creation Illustrated Guide LoRA Practical applications of generative AI rely on LoRA. While managing a large collection of LoRAs, If we discover of similar LoRAs, what kind of LoRA it is! LoRA ? is similaer to Logo LoRA Metadata Output None None LoRA ? what this LoRA is for! aims to apply it to a variety of business Tasks! use LoRA to perform tasks. we can understand

Slide 3

Slide 3 text

Approach: Directly Embedding the LoRA Model Using the Parameters of a Style-Transfer LoRA 2 Input LoRA Weight Parameters Proposed Method Vector that Reflect the Characteristics of LoRA Output ≈ A meaningful sequence of hundreds of dimensions. ≈ tens of millions of parameters. Even without detailed information such as , Metadata Output Examples LoRA models can be represented directly from their weight parameters. Existing Approaches Depend on external information Require metadata or output To represent the LoRA models,

Slide 4

Slide 4 text

Our Method Combines Dimensionality Reduction and Metric Learning. 3 LoRA Weight Parameters Transform LoRA parameters into a compact representation. NN–based Metric Learning LoRA Vector Dimensionality Reduction ・ ・ ・ Represent each LoRA as a sequence of compressed vectors. Embeds LoRA models by applying to their weight parameters. Preprocessing Learning relative similarities among LoRA models. Main component Input Output ・Dimensionality Reduction ・ Metric Learning Overview of the Proposed Method

Slide 5

Slide 5 text

4 ➁ Metric Learning with a Transformer-Based Triplet Network Preprocessing Main Learning Stage Overview of the Proposed Method ➀ Flatten LoRA Parameters from Multiple Layers and Apply PCA-Based Dimensionality Reduction on a Per-Layer Basis

Slide 6

Slide 6 text

Overview of the Proposed Method 5 ➀ Flatten LoRA Parameters from Multiple Layers and Apply PCA-Based Dimensionality Reduction on a Per-Layer Basis ➁ Metric Learning with a Transformer-Based Triplet Network. Preprocessing Main Learning Stage

Slide 7

Slide 7 text

6 LoRA Weight Parameters ・ ・ ・ Text_encoder layer 1 Text_encoder layer 2 ➀ Flatten LoRA Weight and Apply Dimensionality Reduction on a Per-Layer Basis Multiple layers with low-rank matrices. LoRA Down Text_encoder layer 1 Text_encoder layer 2 LoRA param … A compact vector sequence suitable for neural network processing. ➀Flatten each layer individually. To compactly represent LoRA parameters while preserving their characteristics. Incremental PCA Text_encoder layer 3 Text_encoder layer 3 Text_encoder layer1 Text_encoder layer2 Text_encoder layer3 ➁dimensionally reduced on a per-layer basis. ・ ・ ・ Extract layer- specific features Through layer-wise processing, Purpose LoRA Up

Slide 8

Slide 8 text

Overview of the Proposed Method 7 ➀ Flatten LoRA parameters from multiple layers and apply PCA-based dimensionality reduction on a per-layer basis. ➁ Metric Learning with a Transformer-Based Triplet Network Preprocessing Main Learning Stage

Slide 9

Slide 9 text

8 ➁ Metric Learning with a Transformer-Based Triplet Network Assumption Human-perceived similarity between LoRA models is relative rather than absolute. Purpose To learn parameter features based on relative similarities among models. When humans assess the similarity between LoRA models, Van Gogh LoRA LoRA A LoRA B How similar a LoRA model is to a Van Gogh LoRA? Which is more similar to Van Gogh LoRA? Learn LoRA similarities by mimicking human judgments. Goal of Metric Learning Absolute Relative Difficult Easy

Slide 10

Slide 10 text

9 LoRA1 LoRA2 Inputs in Training anchor LoRA3 positive negative Outputs in Training LoRA1 Transformer Encoder Proposed Encoder Per-Layer Proposed Encoder Proposed Encoder LoRA2 LoRA3 Embedding Space After Triplet-Loss Training anchor positive negative Triplet Loss Enables Relative Similarity Learning Learn model similarities in a manner closer to human perception. Pull similar pairs closer Push dissimilar pairs farther apart Triplet Loss Shered weight ➁ Metric Learning with a Transformer-Based Triplet Network To achieve this goal, we adopt a Triplet Network, which processes three inputs using weight-shared encoders. Shered weight MLPs

Slide 11

Slide 11 text

Enable Similarity Computation for LoRA Models without Metadata 10 LoRA Weight Parameters Triplet Metric Learning LoRA Vector Flatten & Dimensionally Reduction ・ ・ ・ Layer-Wise LoRA Embedding Framework Input Output Input Output Retrieve similar LoRA models through vector operations on model parameters. + Transformer Encoder Per-layer MLPs

Slide 12

Slide 12 text

Evaluation of the Proposed Method 11 ➀Evaluation of Learning Method Validity Evaluate how well the embeddings agree with human-labeled triplets. Assess retrieval rankings using human-labeled relevance derived from LoRA output examples. ➁Alignment with Human Similarity Judgments ③LoRA Retrieval Performance Automatic Evaluation Human Evaluation Assess the validity of the proposed framework using prediction accuracy on evaluation triplets. Human Evaluation As parameter-based LoRA retrieval is a new task. Evaluate the proposed components through ablation studies.

Slide 13

Slide 13 text

12 Style-Transfer LoRAs for Stable Diffusion 1.5 Collected from Civitai. Training and Evaluation Data Construction Triplets are constructed based on output similarity. LoRA Dataset Triplets for Training and Evaluation Ground-truth construction based on the assumption that humans judge LoRA similarity from output examples. Sim ≥ 0.6 Anchor LoRA Output Positive LoRA output Negative LoRA output *Thresholds determined from similarity distributions. LoRAs producing similar outputs are treated as similar. Sim ≤ 0.5 Define Training : 549 LoRAs / 464K Triplets Evaluation : 150 LoRAs / 49K Triplets

Slide 14

Slide 14 text

Summary of Experimental Results 13 The proposed embeddings enable stable retrieval of previously unseen LoRA models. The proposed method with positional encoding learns similarity relationships consistent with human perception. The proposed framework effectively learns model similarities from LoRA weight parameters. ➀Evaluation of Learning Method Validity ➁Alignment with Human Similarity Judgments ③LoRA Retrieval Performance

Slide 15

Slide 15 text

Summary of Experimental Results 14 The proposed embeddings enable stable retrieval of previously unseen LoRA models. The proposed method with positional encoding learns similarity relationships consistent with human perception. The proposed framework effectively learns model similarities from LoRA weight parameters. ➀Evaluation of Learning Method Validity ➁Alignment with Human Similarity Judgments ③LoRA Retrieval Performance Why is this Brave New Ideas? A new direction for retrieval: from retrieving multimedia to retrieving multimedia-generating models.

Slide 16

Slide 16 text

Conclusion 15 LoRA Embeddings that Capture Transformation Characteristics from Weight Parameters Key Findings ・Embeddings aligned with human perception. ・Stable retrieval of unseen LoRA models. ➀Layer-wise flattening and Incremental PCA of LoRA parameters. ➁Transformer-based triplet network with layer-importance learning. Metric Learning Preprocessing