メインコンテンツにスキップ
OpenAI

2026年6月30日

GeneBench-Pro の詳細

ベンチマーク、その問題、補足資料を詳しく紹介します。

ケーススタディ

これら10件のケーススタディでは、GeneBench-Pro の代表的な問題を紹介しています。各ケーススタディには、元のプロンプト、データセット、補足資料が含まれています。ベンチマークの概要と主な結果については、発表記事をご覧ください。

注:ファイルのプレビューには、データセット全体からの抜粋が表示されます。


ケーススタディ1

がん体細胞ゲノミクス:構造変異に基づく腫瘍治療のベネフィット・リスク判断

構造変異によって標的の活性化が駆動される腫瘍において、合成 TXR1 標的阻害剤に臨床的有用性があるかどうかを推定します。TXR1、TXR1i、DLR1、およびスターアリルラベルは、合成ベンチマークラベルです。

治療効果と毒性を治療方針の判断材料として解釈する前に、対象サブグループはロングリード解析、発現、腫瘍の品質、薬理ゲノム学的証拠から特定されている必要があります。

モデルに提示された公開済みプロンプト

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

モデルに提供されたファイル

patient_idanalysis_setagesexsitecalendar_periodecogtumor_burdenprior_linesprior_resistancelineage_classtherapy_classassessed16benefit16tox_stop_8wktime_zero_day
MTB0001173.8MS1P220.78731ATXR1i010
MTB0002155.2MS3P112.63701ATXR1i1000MTB0003168.8FS4P200.89121ATXR1i1110MTB0004182.8FS2P224.10100BTXR1i1000MTB0005165.5FS1P317.011ATXR1i1000

レジストリ共変量、治療、16 週時評価、ベネフィット、早期毒性。


ケーススタディ2

機能ゲノミクス:CRISPR 標的検証:lncRNA 転写産物かゲノム遺伝子座か?

見かけ上の lncRNA 依存性が、転写産物特異的なものなのか、それとも近傍遺伝子座や隣接遺伝子の影響によるものなのかを判定します。

転写産物を標的とすることを示すエビデンスは、局所的な DNA 遺伝子座の摂動、近傍遺伝子の抑制、ガイドの入れ替え、GC 毒性、プレート効果に対するコントロールを経てもなお成立している必要があります。

モデルに提示された公開済みプロンプト

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.
  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.
  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;
  • more negative numbers indicate stronger loss of fitness;
  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

モデルに提供されたファイル

guide_idnominal_targetchrcoordstranddist_lnc_tss_bpdist_neighbor_tss_bpguide_gc_frac
g001LINC473chr7100014+14300.624g002LINC473chr7100035-43670.584g003LINC473chr7100051+116560.622g004LINC473chr7100066-59660.617g005LINC473chr7100088+74770.715

ガイドの座標、ターゲット、距離、GC 特徴量。


ケーススタディ3

統計遺伝学:連鎖遺伝子座におけるタンパク質創薬標的の優先順位付け

アッセイ尺度、アリルの向き、ウィナーズカース、連鎖不平衡(LD)、残存する局所的な多面発現を考慮しながら、cis 多変量メンデルランダム化(cis-MVMR)を用いて、近接する 2 つのタンパク質の疾患への直接効果を推定します。

2つのタンパク質は、相関のある遺伝子座を共有しています。解析では、マージナルな関連から、共通のタンパク質尺度で表される、条件付きかつ LD(連鎖不平衡)を考慮した疾患効果へ移行する必要があります。

モデルに提示された公開済みプロンプト

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

モデルに提供されたファイル

snppos_bpeffect_alleleother_allelemafbetasepval
rs20000050000000AC0.422150.0064386683107068080.0032673300912034120.04876727714241972rs20000150010126AC0.057090.0110089933375813010.0069552392087504070.11345916603941006rs20000250020253GT0.090210.0099220147571163190.0056330230270155180.07817048492026045rs20000350030379GT0.483990.0105692156141645730.00322914197402374450.0010638520681901973
rs20000450040506AG0.377030.0070365513782386540.00332975923212698020.034580976884336506

PROTA のスクリーニング段階におけるタンパク質関連サマリー。


ケーススタディ4

臨床ゲノミクス/保因者スクリーニング:CNV と偽遺伝子の較正を踏まえた DRX1 保因者スクリーニングの残余リスク

保因者スクリーニング検査データから、祖先集団別の保因者頻度、スクリーニング陰性後の残余リスク、パートナーの保因者頻度、および罹患受胎産物リスクを推定します。

残余リスクの推定は、偽遺伝子を考慮した保因者判定、創始者ハプロタイプの統合、祖先集団別のアッセイ較正、および検査済みパートナー群から全パートナー名簿への標準化に基づいています。

モデルに提示された公開済みプロンプト

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

モデルに提供されたファイル

sample_idcollectionancestryfamily_history_tier
S_EUR_0001screeningEUR0
S_EUR_0002screeningEUR0
S_EUR_0003screeningEUR0
S_EUR_0004screeningEUR0
S_EUR_0005screeningEUR1

祖先系統とスクリーニング関連情報を含む、スクリーニング名簿上の成人。


ケーススタディ5

シングルセルゲノミクス:アンビエント RNA 補正後の活性化単球 eQTL

シングルセル RNA-seq データからアンビエント RNA と技術的コンタミネーションを除去したうえで、活性化単球における遺伝子型の発現への影響を推定します。

アンビエント RNA は、標的発現と活性化状態の判定に使用されるマーカーパネルの両方に影響するため、eQTL モデルの前に補正を行う必要があります。

モデルに提示された公開済みプロンプト

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

モデルに提供されたファイル

cell_iddonortotal_umiHBBIFI6ISG15LST1CXCL10
D01_C001D011113734835
D01_C002D01110363311210
D01_C003D0111419812639
D01_C004D01125076043217
D01_C005D0110459125115

マーカー遺伝子、コンタミネーションマーカー、および標的遺伝子の細胞ごとの UMI カウント。


ケーススタディ6

構造遺伝学:入れ子状構造変異:発現による裏付けと臨床的関連性

匿名化された逆位様遺伝子座内にある入れ子状の構造的サブハプロタイプに、較正された臨床的関連性と信頼できる発現による裏付けがあるかどうかを推定します。

入れ子状のコピードサージュシグナルは、より広範な逆位の向きによって交絡される可能性があるため、ドサージュ較正、発現による裏付け、臨床モデリングは明確に区別して扱う必要があります。

モデルに提示された公開済みプロンプト

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

モデルに提供されたファイル

sample_idcaseageage_bandsexpc1pc2pc3ancestry_groupclinic_stratumrecruitment_stream
Q00012150.4550_640-1.01514-0.21032-0.08849EURtertiaryclinic
Q00028057.3950_640-1.25987-0.124980.2344EURregionalregistryQ00029168.465_plus00.915980.621770.01891AFRtertiaryclinic
Q00030174.0765_plus10.21125-0.59634-0.08197EAScommunityregistryQ00032182.8265_plus0-1.12034-0.243720.14665EURcommunityclinic

コホート全体の臨床データおよび共変量データ。


ケーススタディ7

制御ゲノミクス:構造変異およびマッピングアーティファクトをマスクした後のクロマチンループ強度の測定

期待接触バックグラウンドから低マッパビリティおよび構造変異由来のアーティファクトを除去したうえで、注目領域における症例・対照間の Hi-C ループ強度差を定量化する。

対象ループは 20 kb 解像度で定義されていますが、低マッパビリティのコンタクトと症例のみに存在する SV ストライプを先にマスクしないと、期待接触モデルが歪みます。

モデルに提示された公開済みプロンプト

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

モデルに提供されたファイル

bin_idchromstartendgc_contentmappabilityre_sites
0chr84000004200000.461990338215725940.97875742147042735
1chr84200004400000.50441242085346770.89010849434983975
2chr84400004600000.432184515849381940.90568792893267123
3chr84600004800000.47331972826812180.93765298406647893
4chr84800005000000.44449560621507480.86825655179818774

ターゲット解像度のビン注釈。


ケーススタディ8

統計遺伝学:創始系統の再構築を用いた多親 QTL マッピング

表現型との関連を検定する前に創始系統由来を再構築し、8 つの創始系統からなる組換え集団で第 1 染色体上の量的形質遺伝子座をマッピングします。

観測可能なマーカーデータは二アレル性ですが、生物学的シグナルは創始系統由来です。したがって、妥当な解析では、創始系統状態を再構築し、マーカーの向きを確認し、QTL をバッチに揃ったノイズピークから分離する必要があります。

モデルに提示された公開済みプロンプト

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata
  • founders.tsv.gz: founder alleles at each marker
  • ril_genotypes.npz: observed RIL genotypes (biallelic)
  • phenotypes.tsv.gz: phenotype and covariates

モデルに提供されたファイル

marker_idchrpos_cM
m2_065259.762431265596575m2_103294.52656615104739m2_107298.18761427503033m2_079272.20130244108847m1_054149.907510212292195

マーカー識別子、染色体、および遺伝地図上の位置。


ケーススタディ9

集団遺伝学:親由来の祖先成分と最近の混合時期

フェーズ決定済みの局所祖先トラクトについて、相互アーティファクトと染色体特異的なラベル反転を修正したうえで、親由来の祖先成分割合と最近の混合時期を推定します。

相互トラクトアーティファクト、染色体内の局所的なラベル反転、またはマップ分母を誤って処理すると、祖先成分割合とパルス時期の両方が変化します。

モデルに提示された公開済みプロンプト

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

モデルに提供されたファイル

chromhapstart_morganend_morganancposteriorlow_complexity_frac
chr1h10.030.505A0.9850.08chr1h10.5050.535B0.620.92chr1h10.5351.478849A0.9850.08chr1h11.5037271.852681B0.9850.08chr1h11.8526812.422373A0.9850.08

座標、祖先系統ラベル、事後確率値、QC 注釈を含む、位相決定済みの局所祖先トラクト。


ケーススタディ10

集団遺伝学:ノイズを含む古代 DNA 時系列データから自然選択を推定する

古代のアリル頻度時系列データから、アリルの向き、方向性誤差、遺伝的浮動、集団サイズの変化を考慮しながら、2 つの半数体遺伝子座のうちどちらがより強い正の自然選択を受けているかを推定します。

ノイズを含む古代のアリル頻度の軌跡は、両方の遺伝子座が同じ派生アリル尺度に揃えられ、提供されたサンプルレベルのシーケンシングエラー値が直接モデル化されるまでは、直接比較できません。

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

モデルに提供されたファイル

generationalt_readstotal_readsseq_errorsample_year
636400.16-4500
1234450.16-42781841550.16-40562438700.16-38333036900.16-3611

遺伝子座 A のリードカウント時系列データ。