跳到主要內容
OpenAI

2026年6月30日

深入了解 GeneBench-Pro

進一步了解這項基準測試、題目及相關支援資料。

實例分析

這 10 項案例研究展示 GeneBench-Pro 中具代表性的題目。每項案例研究均包括原始提示詞、資料集和輔助材料。如要了解基準測試概況和主要發現,請參閱公告網誌

注意:檔案預覽會顯示完整資料集的摘錄。


案例研究 1

體細胞腫瘤學:以結構變異引導的腫瘤治療效益與風險決策

評估一種合成 TXR1 靶向抑制劑,對靶點活化由結構變異驅動的腫瘤是否具有正向臨床效用。TXR1、TXR1i、DLR1 及星號等位基因標籤均為合成基準測試標籤。 

必須先根據長讀長、表達、腫瘤質素及藥物基因組學證據識別目標亞組,才能評估療效和毒性,並據此作出治療決策。

向模型顯示的已發佈提示詞

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

提供給模型的檔案

patient_idanalysis_setagesexsitecalendar_periodecogtumor_burdenprior_linesprior_resistancelineage_classtherapy_classassessed16benefit16tox_stop_8wktime_zero_day
MTB0001173.8MS1P220.78731ATXR1i010
MTB0002155.2MS3P112.63701ATXR1i1000
MTB0003168.8FS4P200.89121ATXR1i1110
MTB0004182.8FS2P224.10100BTXR1i1000
MTB0005165.5FS1P317.011ATXR1i1000

登記資料庫中的協變量、治療、第 16 週評估、療效及早期毒性。


案例研究 2

功能基因組學:CRISPR 靶點驗證:lncRNA 轉錄本還是基因組位點?

判斷看似的 lncRNA 依賴性是轉錄本特異性的,還是由鄰近基因座及鄰近基因效應所驅動。

轉錄本導向證據必須通過多項對照驗證,包括局部 DNA 基因座擾動、鄰近基因抑制、導引序列互換、GC 毒性及培養板效應。

向模型顯示的已發佈提示詞

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.
  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.
  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;
  • more negative numbers indicate stronger loss of fitness;
  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

提供給模型的檔案

guide_idnominal_targetchrcoordstranddist_lnc_tss_bpdist_neighbor_tss_bpguide_gc_frac
g001LINC473chr7100014+14300.624
g002LINC473chr7100035-43670.584
g003LINC473chr7100051+116560.622
g004LINC473chr7100066-59660.617
g005LINC473chr7100088+74770.715

導引序列座標、靶點、距離及 GC 特徵。


案例研究 3

統計遺傳學:在連鎖基因座中優先排序蛋白質藥物靶點

使用順式多變量孟德爾隨機化 (cis-MVMR),估算兩種鄰近蛋白質對疾病的直接效應,同時處理測定尺度、等位基因方向、贏家詛咒、連鎖不平衡 (LD) 及殘餘局部多效性。

這兩種蛋白質共用一個相關基因座。分析必須由邊際關聯轉向條件式、考慮連鎖不平衡 (LD) 的疾病效應,並以統一的蛋白質尺度表示。

向模型提供的已發佈提示詞

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

提供給模型的檔案

snppos_bpeffect_alleleother_allelemafbetasepval
rs20000050000000AC0.422150.0064386683107068080.0032673300912034120.04876727714241972
rs20000150010126AC0.057090.0110089933375813010.0069552392087504070.11345916603941006
rs20000250020253GT0.090210.0099220147571163190.0056330230270155180.07817048492026045
rs20000350030379GT0.483990.0105692156141645730.00322914197402374450.0010638520681901973
rs20000450040506AG0.377030.0070365513782386540.00332975923212698020.034580976884336506

PROTA 的篩選階段蛋白質關聯摘要。


案例研究 4

臨床基因組學/攜帶者篩查:拷貝數變異(CNV)及假基因校準下的 DRX1 攜帶者篩查殘餘風險

根據帶因者篩查檢測資料,估算不同祖源群組的帶因者頻率、陰性篩查後的殘餘風險、伴侶帶因者頻率,以及受影響胚胎的風險。

殘餘風險估算取決於可識別假基因的帶因者判定、創始者單倍型歸併、按祖源群組進行的檢測校準,以及由已檢測伴侶標準化回推至完整伴侶名單。

向模型提供的已發佈提示詞

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

提供給模型的檔案

sample_idcollectionancestryfamily_history_tier
S_EUR_0001screeningEUR0
S_EUR_0002screeningEUR0
S_EUR_0003screeningEUR0
S_EUR_0004screeningEUR0
S_EUR_0005screeningEUR1

附有祖源及篩查背景資料的篩查名冊成年人。


案例研究 5

單細胞基因組學:環境 RNA 校正後的活化單核球 eQTL

從單細胞 RNA-seq 資料中移除環境 RNA 和技術污染後,估算基因型對活化單核細胞基因表達的影響。

環境 RNA 會同時影響目標表達,以及用於判定活化狀態的標記物組合,因此必須在 eQTL 模型之前進行校正。

向模型提供的已發佈提示詞

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

提供給模型的檔案

cell_iddonortotal_umiHBBIFI6ISG15LST1CXCL10
D01_C001D011113734835
D01_C002D01110363311210
D01_C003D0111419812639
D01_C004D01125076043217
D01_C005D0110459125115

每個細胞的標記基因、污染標記及靶基因 UMI 計數。


案例研究 6

結構遺傳學:巢狀結構變異的表達證據及臨床關聯

評估位於未命名類倒位基因座內的巢狀結構亞單倍型,是否具備經校準的臨床關聯和可信的表達證據。

巢狀拷貝劑量訊號可能受較大範圍的倒位方向混淆,因此必須分開處理劑量校準、表達證據和臨床建模。

向模型提供的已發佈提示詞

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

提供給模型的檔案

sample_idcaseageage_bandsexpc1pc2pc3ancestry_groupclinic_stratumrecruitment_stream
Q00012150.4550_640-1.01514-0.21032-0.08849EURtertiaryclinic
Q00028057.3950_640-1.25987-0.124980.2344EURregionalregistry
Q00029168.465_plus00.915980.621770.01891AFRtertiaryclinic
Q00030174.0765_plus10.21125-0.59634-0.08197EAScommunityregistry
Q00032182.8265_plus0-1.12034-0.243720.14665EURcommunityclinic

完整隊列的臨床及協變量資料。


案例研究 7

調控基因組學:遮罩結構變異和比對偽影後測量染色質環路強度

從預期接觸背景中移除低可比對性及結構變異偽影後,量化局部病例對照 Hi-C 環路強度的差異。

目標環路以 20 kb 解像度界定,但必須先遮罩低可比對性接觸和病例特有的 SV 條帶,否則預期接觸模型會失真。

向模型提供的已發佈提示詞

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

提供給模型的檔案

bin_idchromstartendgc_contentmappabilityre_sites
0chr84000004200000.461990338215725940.97875742147042735
1chr84200004400000.50441242085346770.89010849434983975
2chr84400004600000.432184515849381940.90568792893267123
3chr84600004800000.47331972826812180.93765298406647893
4chr84800005000000.44449560621507480.86825655179818774

目標解析度分箱註釋。


案例研究 8

統計遺傳學:多親本 QTL 映射與創始者重建

在八個創始親本的重組族群中,先重建創始親本祖源,再測試表型關聯,以定位第一號染色體上的數量性狀基因座。

可見的標記資料是雙等位基因的,但生物學訊號則是創始者祖源。因此,一項站得住腳的分析必須重建創始者狀態、檢查標記方向,並將 QTL(數量性狀基因座)與一個與批次效應相符的干擾峰區分開來。

向模型提供的已發佈提示詞

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata
  • founders.tsv.gz: founder alleles at each marker
  • ril_genotypes.npz: observed RIL genotypes (biallelic)
  • phenotypes.tsv.gz: phenotype and covariates

提供給模型的檔案

marker_idchrpos_cM
m2_065259.762431265596575
m2_103294.52656615104739
m2_107298.18761427503033
m2_079272.20130244108847
m1_054149.907510212292195

標記識別碼、染色體及遺傳圖譜位置。


案例研究 9

群體遺傳學:親本特異性祖源與近期混合時間

修正互易偽影及特定染色體的標籤反轉後,根據已定相的局部祖源片段,推斷父母各自的祖源比例及近期基因混合時間。

如果互易片段偽影、染色體局部標籤反轉或圖譜分母處理不當,祖源比例和脈衝時間都會改變。

向模型提供的已發佈提示詞

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

提供給模型的檔案

chromhapstart_morganend_morganancposteriorlow_complexity_frac
chr1h10.030.505A0.9850.08
chr1h10.5050.535B0.620.92
chr1h10.5351.478849A0.9850.08
chr1h11.5037271.852681B0.9850.08
chr1h11.8526812.422373A0.9850.08

已定相的局部祖源區段,包含座標、祖源標籤、後驗值及質控註釋。


案例研究 10

群體遺傳學:從帶雜訊的古代 DNA 時間序列估算自然選擇

根據古代等位基因頻率的時間序列,推斷兩個單倍體基因座中哪一個受到較強的正向選擇,同時考慮等位基因方向、方向性誤差、遺傳漂變及族群大小變化。

在把兩個基因座置於相同的衍生等位基因尺度,並直接對所提供的樣本層級測序誤差值建模前,無法直接比較含有雜訊的古代軌跡。

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

提供給模型的檔案

generationalt_readstotal_readsseq_errorsample_year
636400.16-4500
1234450.16-4278
1841550.16-4056
2438700.16-3833
3036900.16-3611

基因座 A 的讀段計數時間序列。