{
  "id": 687757,
  "title": "33rd Place Solution",
  "url": "/competitions/stanford-rna-3d-folding-2/discussion/687757",
  "author_name": "t fuku",
  "post_date": "2026-04-03T15:46:51.039000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<h1>Stanford RNA 3D Folding Part 2 — 33rd Place Solution (Silver Medal)</h1>\n<p><strong>Final Submission: ex77 Adaptive Diversity Selection</strong></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Best Private LB</strong></td>\n<td>0.447</td>\n</tr>\n<tr>\n<td><strong>Best Public LB</strong></td>\n<td>0.433</td>\n</tr>\n<tr>\n<td><strong>Private LB Rank</strong></td>\n<td>33rd / 1,877 (Silver Medal)</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Solution: Overall Approach</h2>\n<p>Our solution <strong>ex77 Adaptive Diversity Selection</strong> is an ensemble pipeline that combines 3 structure prediction methods and <strong>adaptively determines the optimal slot allocation per target</strong>.</p>\n<ul>\n<li><strong>Phase 1: TBM (Template-Based Modeling)</strong> -- Template matching based on sequence similarity</li>\n<li><strong>Phase 2a: DRfold2</strong> -- Deep learning prediction for short-chain RNA</li>\n<li><strong>Phase 2b: Protenix</strong> -- Diffusion model-based structure prediction</li>\n<li><strong>Phase 3: Adaptive Selection</strong> -- Diversity maximization via Kabsch RMSD + FPS</li>\n</ul>\n<h3>Pipeline Diagram</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2Fdf211bdd17120f7e972174d9020239d3%2Fpipeline.png?generation=1775231362234073&amp;alt=media\" alt=\"pipeline\">\n<em>Figure 1: Solution Pipeline -- 3-method ensemble + adaptive diversity selection</em></p>\n<h3>Phase 1: TBM (Template-Based Modeling)</h3>\n<p>Template-based modeling is the <strong>core component</strong> of this solution. It uses known structures from training data as templates and transfers 3D coordinates based on sequence alignment.</p>\n<h4>Sequence Alignment</h4>\n<p>Global alignment using BioPython's <code>PairwiseAligner</code>:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>match_score</td>\n<td>2</td>\n<td>Match bonus</td>\n</tr>\n<tr>\n<td>mismatch_score</td>\n<td>-1.5</td>\n<td>Mismatch penalty</td>\n</tr>\n<tr>\n<td>open_gap_score</td>\n<td>-8</td>\n<td>Gap opening</td>\n</tr>\n<tr>\n<td>extend_gap_score</td>\n<td>-0.4</td>\n<td>Gap extension</td>\n</tr>\n</tbody>\n</table>\n<h4>Template Adaptation</h4>\n<p>The <code>adapt_template_to_query</code> function processes alignment results:</p>\n<ul>\n<li><strong>Match positions</strong>: Directly transfer template 3D coordinates</li>\n<li><strong>Gap positions</strong>: Linear interpolation from flanking known coordinates</li>\n<li><strong>Terminal gaps</strong>: Extrapolation at 3.0A intervals from nearest known coordinates</li>\n</ul>\n<h4>Template Pool</h4>\n<p>Integrated <code>train_sequences</code> + <code>validation_sequences</code> (~5,700 sequences). Generates <strong>top_n=30</strong> candidates per target.</p>\n<h3>Phase 2a: DRfold2</h3>\n<p><strong>DRfold2</strong> is a deep learning structure prediction model specialized for short-chain RNA (100nt or less).</p>\n<ul>\n<li>Target: Only sequences with length &lt;= 100nt</li>\n<li>Model: cfg_97 configuration</li>\n<li>Time limit: 2 hours</li>\n<li>Output: Extract C1' atom coordinates from PDB files</li>\n</ul>\n<p>DRfold2 predictions are added to the diversity pool and utilized in Phase 3 FPS selection.</p>\n<h3>Phase 2b: Protenix</h3>\n<p><strong>Protenix</strong> is an AlphaFold-based diffusion model structure prediction tool that runs inference on all targets.</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>N_SAMPLE</td>\n<td>5</td>\n</tr>\n<tr>\n<td>SEED</td>\n<td>42</td>\n</tr>\n<tr>\n<td>MAX_SEQ_LEN</td>\n<td>512 (absolute limit)</td>\n</tr>\n<tr>\n<td>CHUNK_OVERLAP</td>\n<td>128</td>\n</tr>\n<tr>\n<td>USE_RNA_MSA</td>\n<td>true</td>\n</tr>\n<tr>\n<td>USE_MSA</td>\n<td>false</td>\n</tr>\n<tr>\n<td>USE_TEMPLATE</td>\n<td>false</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p><strong>MAX_SEQ_LEN=512 is the absolute limit</strong> -- Setting it to 600 causes OOM on Kaggle T4 GPU (16GB VRAM), resulting in scoring failure. Sequences exceeding 512 are handled via chunking.</p>\n</blockquote>\n<h3>Phase 3: Adaptive Diversity Selection</h3>\n<p>The <strong>key innovation</strong> of this solution is the mechanism that adaptively determines TBM slot count per target.</p>\n<ul>\n<li><strong>top pct_id &gt;= 50%</strong>: TBM 3 slots + diversity pool 2 slots</li>\n<li><strong>top pct_id &lt; 50%</strong>: TBM 0 slots + diversity pool 5 slots</li>\n</ul>\n<p>The diversity pool selects the most diverse structures from all DRfold2 + Protenix candidates using <strong>Farthest Point Sampling (FPS)</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2F66e050ce8ef6227581fb667fc8e1cfe1%2Fadaptive_slot.png?generation=1775231394035642&amp;alt=media\" alt=\"adaptive\">\n<em>Figure 2: Adaptive slot allocation logic</em></p>\n<h4>Farthest Point Sampling (FPS)</h4>\n<p>Algorithm to maximize diversity:</p>\n<ol>\n<li>Compute <strong>Kabsch RMSD matrix</strong> between all candidates</li>\n<li>Select the candidate with the largest average RMSD first</li>\n<li>Select the candidate with the maximum minimum RMSD to already-selected candidates</li>\n<li>Repeat until the required number of slots is filled</li>\n</ol>\n<p>This maximizes the structural space covered by the 5 predictions, improving the Best-of-5 TM-score.</p>\n<h3>Chunking Strategy</h3>\n<p>For long RNA sequences exceeding MAX_SEQ_LEN=512 (e.g., 9ZCC=1460nt, 9MME=4168nt), we use <strong>overlapping chunk splitting</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2F548f4c8b1e7e001d8d3017b27e37037b%2Fchunking.png?generation=1775231409404538&amp;alt=media\" alt=\"chunking\">\n<em>Figure 3: Chunking strategy for long RNA sequences</em></p>\n<h4>Stitching (Assembly)</h4>\n<p>Procedure to combine Protenix outputs from each chunk into full-length coordinates:</p>\n<ol>\n<li>Compute optimal rotation/translation via Kabsch alignment (SVD decomposition) on overlap regions</li>\n<li>Align subsequent chunk coordinates to the previous chunk</li>\n<li>Use <strong>linear blending</strong> (weighted average) on overlap regions for smooth connection</li>\n</ol>\n<h3>RNA Physical Constraints</h3>\n<p>The <code>adaptive_rna_constraints</code> function applies physical plausibility to each structure:</p>\n<table>\n<thead>\n<tr>\n<th>Constraint</th>\n<th>Parameter</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Bond distance</td>\n<td>5.95 A</td>\n<td>Ideal distance between adjacent C1' atoms</td>\n</tr>\n<tr>\n<td>Next-nearest distance</td>\n<td>10.2 A</td>\n<td>Ideal distance between C1' atoms 2 residues apart</td>\n</tr>\n<tr>\n<td>Laplacian smoothing</td>\n<td>0.06</td>\n<td>Local coordinate smoothing</td>\n</tr>\n<tr>\n<td>Clash avoidance</td>\n<td>3.2 A</td>\n<td>Minimum distance between non-adjacent residues</td>\n</tr>\n<tr>\n<td>Passes</td>\n<td>2</td>\n<td>Number of constraint application iterations</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 3435038,
      "postDate": "2026-04-03T15:46:51.040Z",
      "content": "<h1>Stanford RNA 3D Folding Part 2 — 33rd Place Solution (Silver Medal)</h1>\n<p><strong>Final Submission: ex77 Adaptive Diversity Selection</strong></p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Best Private LB</strong></td>\n<td>0.447</td>\n</tr>\n<tr>\n<td><strong>Best Public LB</strong></td>\n<td>0.433</td>\n</tr>\n<tr>\n<td><strong>Private LB Rank</strong></td>\n<td>33rd / 1,877 (Silver Medal)</td>\n</tr>\n</tbody>\n</table>\n<hr>\n<h2>Solution: Overall Approach</h2>\n<p>Our solution <strong>ex77 Adaptive Diversity Selection</strong> is an ensemble pipeline that combines 3 structure prediction methods and <strong>adaptively determines the optimal slot allocation per target</strong>.</p>\n<ul>\n<li><strong>Phase 1: TBM (Template-Based Modeling)</strong> -- Template matching based on sequence similarity</li>\n<li><strong>Phase 2a: DRfold2</strong> -- Deep learning prediction for short-chain RNA</li>\n<li><strong>Phase 2b: Protenix</strong> -- Diffusion model-based structure prediction</li>\n<li><strong>Phase 3: Adaptive Selection</strong> -- Diversity maximization via Kabsch RMSD + FPS</li>\n</ul>\n<h3>Pipeline Diagram</h3>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2Fdf211bdd17120f7e972174d9020239d3%2Fpipeline.png?generation=1775231362234073&amp;alt=media\" alt=\"pipeline\">\n<em>Figure 1: Solution Pipeline -- 3-method ensemble + adaptive diversity selection</em></p>\n<h3>Phase 1: TBM (Template-Based Modeling)</h3>\n<p>Template-based modeling is the <strong>core component</strong> of this solution. It uses known structures from training data as templates and transfers 3D coordinates based on sequence alignment.</p>\n<h4>Sequence Alignment</h4>\n<p>Global alignment using BioPython's <code>PairwiseAligner</code>:</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n<th>Note</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>match_score</td>\n<td>2</td>\n<td>Match bonus</td>\n</tr>\n<tr>\n<td>mismatch_score</td>\n<td>-1.5</td>\n<td>Mismatch penalty</td>\n</tr>\n<tr>\n<td>open_gap_score</td>\n<td>-8</td>\n<td>Gap opening</td>\n</tr>\n<tr>\n<td>extend_gap_score</td>\n<td>-0.4</td>\n<td>Gap extension</td>\n</tr>\n</tbody>\n</table>\n<h4>Template Adaptation</h4>\n<p>The <code>adapt_template_to_query</code> function processes alignment results:</p>\n<ul>\n<li><strong>Match positions</strong>: Directly transfer template 3D coordinates</li>\n<li><strong>Gap positions</strong>: Linear interpolation from flanking known coordinates</li>\n<li><strong>Terminal gaps</strong>: Extrapolation at 3.0A intervals from nearest known coordinates</li>\n</ul>\n<h4>Template Pool</h4>\n<p>Integrated <code>train_sequences</code> + <code>validation_sequences</code> (~5,700 sequences). Generates <strong>top_n=30</strong> candidates per target.</p>\n<h3>Phase 2a: DRfold2</h3>\n<p><strong>DRfold2</strong> is a deep learning structure prediction model specialized for short-chain RNA (100nt or less).</p>\n<ul>\n<li>Target: Only sequences with length &lt;= 100nt</li>\n<li>Model: cfg_97 configuration</li>\n<li>Time limit: 2 hours</li>\n<li>Output: Extract C1' atom coordinates from PDB files</li>\n</ul>\n<p>DRfold2 predictions are added to the diversity pool and utilized in Phase 3 FPS selection.</p>\n<h3>Phase 2b: Protenix</h3>\n<p><strong>Protenix</strong> is an AlphaFold-based diffusion model structure prediction tool that runs inference on all targets.</p>\n<table>\n<thead>\n<tr>\n<th>Parameter</th>\n<th>Value</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>N_SAMPLE</td>\n<td>5</td>\n</tr>\n<tr>\n<td>SEED</td>\n<td>42</td>\n</tr>\n<tr>\n<td>MAX_SEQ_LEN</td>\n<td>512 (absolute limit)</td>\n</tr>\n<tr>\n<td>CHUNK_OVERLAP</td>\n<td>128</td>\n</tr>\n<tr>\n<td>USE_RNA_MSA</td>\n<td>true</td>\n</tr>\n<tr>\n<td>USE_MSA</td>\n<td>false</td>\n</tr>\n<tr>\n<td>USE_TEMPLATE</td>\n<td>false</td>\n</tr>\n</tbody>\n</table>\n<blockquote>\n  <p><strong>MAX_SEQ_LEN=512 is the absolute limit</strong> -- Setting it to 600 causes OOM on Kaggle T4 GPU (16GB VRAM), resulting in scoring failure. Sequences exceeding 512 are handled via chunking.</p>\n</blockquote>\n<h3>Phase 3: Adaptive Diversity Selection</h3>\n<p>The <strong>key innovation</strong> of this solution is the mechanism that adaptively determines TBM slot count per target.</p>\n<ul>\n<li><strong>top pct_id &gt;= 50%</strong>: TBM 3 slots + diversity pool 2 slots</li>\n<li><strong>top pct_id &lt; 50%</strong>: TBM 0 slots + diversity pool 5 slots</li>\n</ul>\n<p>The diversity pool selects the most diverse structures from all DRfold2 + Protenix candidates using <strong>Farthest Point Sampling (FPS)</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2F66e050ce8ef6227581fb667fc8e1cfe1%2Fadaptive_slot.png?generation=1775231394035642&amp;alt=media\" alt=\"adaptive\">\n<em>Figure 2: Adaptive slot allocation logic</em></p>\n<h4>Farthest Point Sampling (FPS)</h4>\n<p>Algorithm to maximize diversity:</p>\n<ol>\n<li>Compute <strong>Kabsch RMSD matrix</strong> between all candidates</li>\n<li>Select the candidate with the largest average RMSD first</li>\n<li>Select the candidate with the maximum minimum RMSD to already-selected candidates</li>\n<li>Repeat until the required number of slots is filled</li>\n</ol>\n<p>This maximizes the structural space covered by the 5 predictions, improving the Best-of-5 TM-score.</p>\n<h3>Chunking Strategy</h3>\n<p>For long RNA sequences exceeding MAX_SEQ_LEN=512 (e.g., 9ZCC=1460nt, 9MME=4168nt), we use <strong>overlapping chunk splitting</strong>.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2F548f4c8b1e7e001d8d3017b27e37037b%2Fchunking.png?generation=1775231409404538&amp;alt=media\" alt=\"chunking\">\n<em>Figure 3: Chunking strategy for long RNA sequences</em></p>\n<h4>Stitching (Assembly)</h4>\n<p>Procedure to combine Protenix outputs from each chunk into full-length coordinates:</p>\n<ol>\n<li>Compute optimal rotation/translation via Kabsch alignment (SVD decomposition) on overlap regions</li>\n<li>Align subsequent chunk coordinates to the previous chunk</li>\n<li>Use <strong>linear blending</strong> (weighted average) on overlap regions for smooth connection</li>\n</ol>\n<h3>RNA Physical Constraints</h3>\n<p>The <code>adaptive_rna_constraints</code> function applies physical plausibility to each structure:</p>\n<table>\n<thead>\n<tr>\n<th>Constraint</th>\n<th>Parameter</th>\n<th>Description</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Bond distance</td>\n<td>5.95 A</td>\n<td>Ideal distance between adjacent C1' atoms</td>\n</tr>\n<tr>\n<td>Next-nearest distance</td>\n<td>10.2 A</td>\n<td>Ideal distance between C1' atoms 2 residues apart</td>\n</tr>\n<tr>\n<td>Laplacian smoothing</td>\n<td>0.06</td>\n<td>Local coordinate smoothing</td>\n</tr>\n<tr>\n<td>Clash avoidance</td>\n<td>3.2 A</td>\n<td>Minimum distance between non-adjacent residues</td>\n</tr>\n<tr>\n<td>Passes</td>\n<td>2</td>\n<td>Number of constraint application iterations</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "# Stanford RNA 3D Folding Part 2 — 33rd Place Solution (Silver Medal)\n\n**Final Submission: ex77 Adaptive Diversity Selection**\n\n| | |\n|---|---|\n| **Best Private LB** | 0.447 |\n| **Best Public LB** | 0.433 |\n| **Private LB Rank** | 33rd / 1,877 (Silver Medal) |\n\n---\n\n## Solution: Overall Approach\n\nOur solution **ex77 Adaptive Diversity Selection** is an ensemble pipeline that combines 3 structure prediction methods and **adaptively determines the optimal slot allocation per target**.\n\n- **Phase 1: TBM (Template-Based Modeling)** -- Template matching based on sequence similarity\n- **Phase 2a: DRfold2** -- Deep learning prediction for short-chain RNA\n- **Phase 2b: Protenix** -- Diffusion model-based structure prediction\n- **Phase 3: Adaptive Selection** -- Diversity maximization via Kabsch RMSD + FPS\n\n### Pipeline Diagram\n\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2Fdf211bdd17120f7e972174d9020239d3%2Fpipeline.png?generation=1775231362234073&alt=media)\n*Figure 1: Solution Pipeline -- 3-method ensemble + adaptive diversity selection*\n\n### Phase 1: TBM (Template-Based Modeling)\n\nTemplate-based modeling is the **core component** of this solution. It uses known structures from training data as templates and transfers 3D coordinates based on sequence alignment.\n\n#### Sequence Alignment\n\nGlobal alignment using BioPython's `PairwiseAligner`:\n\n| Parameter | Value | Note |\n|-----------|-------|------|\n| match_score | 2 | Match bonus |\n| mismatch_score | -1.5 | Mismatch penalty |\n| open_gap_score | -8 | Gap opening |\n| extend_gap_score | -0.4 | Gap extension |\n\n#### Template Adaptation\n\nThe `adapt_template_to_query` function processes alignment results:\n\n- **Match positions**: Directly transfer template 3D coordinates\n- **Gap positions**: Linear interpolation from flanking known coordinates\n- **Terminal gaps**: Extrapolation at 3.0A intervals from nearest known coordinates\n\n#### Template Pool\n\nIntegrated `train_sequences` + `validation_sequences` (~5,700 sequences). Generates **top_n=30** candidates per target.\n\n### Phase 2a: DRfold2\n\n**DRfold2** is a deep learning structure prediction model specialized for short-chain RNA (100nt or less).\n\n- Target: Only sequences with length <= 100nt\n- Model: cfg_97 configuration\n- Time limit: 2 hours\n- Output: Extract C1' atom coordinates from PDB files\n\nDRfold2 predictions are added to the diversity pool and utilized in Phase 3 FPS selection.\n\n### Phase 2b: Protenix\n\n**Protenix** is an AlphaFold-based diffusion model structure prediction tool that runs inference on all targets.\n\n| Parameter | Value |\n|-----------|-------|\n| N_SAMPLE | 5 |\n| SEED | 42 |\n| MAX_SEQ_LEN | 512 (absolute limit) |\n| CHUNK_OVERLAP | 128 |\n| USE_RNA_MSA | true |\n| USE_MSA | false |\n| USE_TEMPLATE | false |\n\n> **MAX_SEQ_LEN=512 is the absolute limit** -- Setting it to 600 causes OOM on Kaggle T4 GPU (16GB VRAM), resulting in scoring failure. Sequences exceeding 512 are handled via chunking.\n\n### Phase 3: Adaptive Diversity Selection\n\nThe **key innovation** of this solution is the mechanism that adaptively determines TBM slot count per target.\n\n- **top pct_id >= 50%**: TBM 3 slots + diversity pool 2 slots\n- **top pct_id < 50%**: TBM 0 slots + diversity pool 5 slots\n\nThe diversity pool selects the most diverse structures from all DRfold2 + Protenix candidates using **Farthest Point Sampling (FPS)**.\n\n![adaptive](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2F66e050ce8ef6227581fb667fc8e1cfe1%2Fadaptive_slot.png?generation=1775231394035642&alt=media)\n*Figure 2: Adaptive slot allocation logic*\n\n#### Farthest Point Sampling (FPS)\n\nAlgorithm to maximize diversity:\n\n1. Compute **Kabsch RMSD matrix** between all candidates\n2. Select the candidate with the largest average RMSD first\n3. Select the candidate with the maximum minimum RMSD to already-selected candidates\n4. Repeat until the required number of slots is filled\n\nThis maximizes the structural space covered by the 5 predictions, improving the Best-of-5 TM-score.\n\n### Chunking Strategy\n\nFor long RNA sequences exceeding MAX_SEQ_LEN=512 (e.g., 9ZCC=1460nt, 9MME=4168nt), we use **overlapping chunk splitting**.\n\n![chunking](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2F548f4c8b1e7e001d8d3017b27e37037b%2Fchunking.png?generation=1775231409404538&alt=media)\n*Figure 3: Chunking strategy for long RNA sequences*\n\n#### Stitching (Assembly)\n\nProcedure to combine Protenix outputs from each chunk into full-length coordinates:\n\n1. Compute optimal rotation/translation via Kabsch alignment (SVD decomposition) on overlap regions\n2. Align subsequent chunk coordinates to the previous chunk\n3. Use **linear blending** (weighted average) on overlap regions for smooth connection\n\n### RNA Physical Constraints\n\nThe `adaptive_rna_constraints` function applies physical plausibility to each structure:\n\n| Constraint | Parameter | Description |\n|-----------|-----------|-------------|\n| Bond distance | 5.95 A | Ideal distance between adjacent C1' atoms |\n| Next-nearest distance | 10.2 A | Ideal distance between C1' atoms 2 residues apart |\n| Laplacian smoothing | 0.06 | Local coordinate smoothing |\n| Clash avoidance | 3.2 A | Minimum distance between non-adjacent residues |\n| Passes | 2 | Number of constraint application iterations |\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3435038": "# Stanford RNA 3D Folding Part 2 — 33rd Place Solution (Silver Medal)\n\n**Final Submission: ex77 Adaptive Diversity Selection**\n\n| | |\n|---|---|\n| **Best Private LB** | 0.447 |\n| **Best Public LB** | 0.433 |\n| **Private LB Rank** | 33rd / 1,877 (Silver Medal) |\n\n---\n\n## Solution: Overall Approach\n\nOur solution **ex77 Adaptive Diversity Selection** is an ensemble pipeline that combines 3 structure prediction methods and **adaptively determines the optimal slot allocation per target**.\n\n- **Phase 1: TBM (Template-Based Modeling)** -- Template matching based on sequence similarity\n- **Phase 2a: DRfold2** -- Deep learning prediction for short-chain RNA\n- **Phase 2b: Protenix** -- Diffusion model-based structure prediction\n- **Phase 3: Adaptive Selection** -- Diversity maximization via Kabsch RMSD + FPS\n\n### Pipeline Diagram\n\n![pipeline](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2Fdf211bdd17120f7e972174d9020239d3%2Fpipeline.png?generation=1775231362234073&alt=media)\n*Figure 1: Solution Pipeline -- 3-method ensemble + adaptive diversity selection*\n\n### Phase 1: TBM (Template-Based Modeling)\n\nTemplate-based modeling is the **core component** of this solution. It uses known structures from training data as templates and transfers 3D coordinates based on sequence alignment.\n\n#### Sequence Alignment\n\nGlobal alignment using BioPython's `PairwiseAligner`:\n\n| Parameter | Value | Note |\n|-----------|-------|------|\n| match_score | 2 | Match bonus |\n| mismatch_score | -1.5 | Mismatch penalty |\n| open_gap_score | -8 | Gap opening |\n| extend_gap_score | -0.4 | Gap extension |\n\n#### Template Adaptation\n\nThe `adapt_template_to_query` function processes alignment results:\n\n- **Match positions**: Directly transfer template 3D coordinates\n- **Gap positions**: Linear interpolation from flanking known coordinates\n- **Terminal gaps**: Extrapolation at 3.0A intervals from nearest known coordinates\n\n#### Template Pool\n\nIntegrated `train_sequences` + `validation_sequences` (~5,700 sequences). Generates **top_n=30** candidates per target.\n\n### Phase 2a: DRfold2\n\n**DRfold2** is a deep learning structure prediction model specialized for short-chain RNA (100nt or less).\n\n- Target: Only sequences with length <= 100nt\n- Model: cfg_97 configuration\n- Time limit: 2 hours\n- Output: Extract C1' atom coordinates from PDB files\n\nDRfold2 predictions are added to the diversity pool and utilized in Phase 3 FPS selection.\n\n### Phase 2b: Protenix\n\n**Protenix** is an AlphaFold-based diffusion model structure prediction tool that runs inference on all targets.\n\n| Parameter | Value |\n|-----------|-------|\n| N_SAMPLE | 5 |\n| SEED | 42 |\n| MAX_SEQ_LEN | 512 (absolute limit) |\n| CHUNK_OVERLAP | 128 |\n| USE_RNA_MSA | true |\n| USE_MSA | false |\n| USE_TEMPLATE | false |\n\n> **MAX_SEQ_LEN=512 is the absolute limit** -- Setting it to 600 causes OOM on Kaggle T4 GPU (16GB VRAM), resulting in scoring failure. Sequences exceeding 512 are handled via chunking.\n\n### Phase 3: Adaptive Diversity Selection\n\nThe **key innovation** of this solution is the mechanism that adaptively determines TBM slot count per target.\n\n- **top pct_id >= 50%**: TBM 3 slots + diversity pool 2 slots\n- **top pct_id < 50%**: TBM 0 slots + diversity pool 5 slots\n\nThe diversity pool selects the most diverse structures from all DRfold2 + Protenix candidates using **Farthest Point Sampling (FPS)**.\n\n![adaptive](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2F66e050ce8ef6227581fb667fc8e1cfe1%2Fadaptive_slot.png?generation=1775231394035642&alt=media)\n*Figure 2: Adaptive slot allocation logic*\n\n#### Farthest Point Sampling (FPS)\n\nAlgorithm to maximize diversity:\n\n1. Compute **Kabsch RMSD matrix** between all candidates\n2. Select the candidate with the largest average RMSD first\n3. Select the candidate with the maximum minimum RMSD to already-selected candidates\n4. Repeat until the required number of slots is filled\n\nThis maximizes the structural space covered by the 5 predictions, improving the Best-of-5 TM-score.\n\n### Chunking Strategy\n\nFor long RNA sequences exceeding MAX_SEQ_LEN=512 (e.g., 9ZCC=1460nt, 9MME=4168nt), we use **overlapping chunk splitting**.\n\n![chunking](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3790528%2F548f4c8b1e7e001d8d3017b27e37037b%2Fchunking.png?generation=1775231409404538&alt=media)\n*Figure 3: Chunking strategy for long RNA sequences*\n\n#### Stitching (Assembly)\n\nProcedure to combine Protenix outputs from each chunk into full-length coordinates:\n\n1. Compute optimal rotation/translation via Kabsch alignment (SVD decomposition) on overlap regions\n2. Align subsequent chunk coordinates to the previous chunk\n3. Use **linear blending** (weighted average) on overlap regions for smooth connection\n\n### RNA Physical Constraints\n\nThe `adaptive_rna_constraints` function applies physical plausibility to each structure:\n\n| Constraint | Parameter | Description |\n|-----------|-----------|-------------|\n| Bond distance | 5.95 A | Ideal distance between adjacent C1' atoms |\n| Next-nearest distance | 10.2 A | Ideal distance between C1' atoms 2 residues apart |\n| Laplacian smoothing | 0.06 | Local coordinate smoothing |\n| Clash avoidance | 3.2 A | Minimum distance between non-adjacent residues |\n| Passes | 2 | Number of constraint application iterations |\n"
  }
}