Vai al contenuto principale
OpenAI

30 giugno 2026

All'interno di GeneBench-Pro

Uno sguardo più approfondito al benchmark, alle sue domande e ai materiali di supporto.

Casi di studio

Dieci casi di studio illustrano domande rappresentative tratte da GeneBench-Pro. Ogni caso di studio include il prompt originale, i dataset e i materiali di supporto. Per una panoramica del benchmark e dei risultati principali, consulta il blog di annuncio.

Nota: le anteprime dei file mostrano estratti dai dataset completi.


Caso di studio 1

Oncologia somatica: decisione sul rapporto beneficio-rischio della terapia antitumorale guidata da varianti strutturali

Stima se un inibitore sintetico mirato a TXR1 abbia un'utilità clinica nei tumori in cui l'attivazione del target è determinata da una variante strutturale. TXR1, TXR1i, DLR1 e le etichette star-allele sono etichette di benchmark sintetiche. 

Il sottogruppo target deve essere identificato sulla base delle evidenze derivanti dalle letture lunghe, dall'espressione, dalla qualità tumorale e dalla farmacogenomica prima che beneficio e tossicità possano essere interpretati ai fini della decisione terapeutica.

Prompt rilasciato mostrato al modello

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

File forniti al modello


Caso di studio 2

Genomica funzionale: validazione del target con CRISPR: trascritto di lncRNA o locus genomico?

Determina se un'apparente dipendenza da lncRNA sia specifica del trascritto o dovuta agli effetti del locus adiacente e dei geni vicini.

Le evidenze dirette sul trascritto devono superare i controlli relativi alla perturbazione locale del locus del DNA, alla repressione dei geni vicini, agli scambi di guide, alla tossicità da GC e agli effetti di piastra.

Prompt rilasciato mostrato al modello

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.

  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.

  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;

  • more negative numbers indicate stronger loss of fitness;

  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

File forniti al modello


Caso di studio 3

Genetica statistica: prioritizzazione dei bersagli farmacologici proteici in un locus genetico associato

Stima gli effetti diretti sulla malattia per due proteine vicine mediante randomizzazione mendeliana multivariabile cis (cis-MVMR), tenendo conto della scala del saggio, l’orientamento allelico, la maledizione del vincitore, il disequilibrio di linkage (LD) e la pleiotropia locale residua.

Le due proteine condividono un locus genetico associato. L'analisi deve passare dalle associazioni marginali agli effetti condizionali della malattia, corretti per il disequilibrio di linkage (LD), su una scala proteica comune.

Prompt rilasciato mostrato al modello

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

File forniti al modello


Caso di studio 4

Genomica clinica / screening dei portatori: rischio residuo dello screening dei portatori per DRX1 con calibrazione per CNV e pseudogeni

Stima le frequenze specifiche dei portatori per ascendenza, il rischio residuo dopo uno screening negativo, la frequenza dei portatori nel partner e il rischio che il concepito sia affetto, basandoti sui dati del test di screening dei portatori.

La stima del rischio residuo dipende dall'identificazione dello stato di portatore corretta per gli pseudogeni, dal collasso degli aplotipi fondatori, dalla calibrazione del saggio specifica per l'ascendenza e dalla standardizzazione dai partner testati all'intero registro dei partner.

Prompt rilasciato mostrato al modello

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

File forniti al modello


Caso di studio 5

Genomica a singola cellula: eQTL dei monociti attivati dopo correzione dell’RNA ambientale

Stima un effetto del genotipo sull'espressione dei monociti attivati dopo aver rimosso l'RNA ambientale e la contaminazione tecnica dai dati di RNA-seq a singola cellula.

L’RNA ambientale influisce sia sull’espressione del target sia sul pannello di marcatori utilizzato per determinare lo stato di attivazione, quindi la correzione deve avvenire prima del modello eQTL.

Prompt rilasciato mostrato al modello

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

File forniti al modello


Caso di studio 6

Genetica strutturale: variante strutturale annidata: supporto dell’espressione genica e associazione clinica

Stimare se un subaplotipo strutturale annidato all’interno di un locus anonimo simile a un’inversione abbia un’associazione clinica calibrata e un supporto credibile a livello di espressione.

Un segnale annidato di dosaggio delle copie può essere soggetto a confondimento da parte dell’orientamento dell’inversione più ampia, pertanto la calibrazione del dosaggio, il supporto dell’espressione e la modellizzazione clinica devono rimanere distinti.

Prompt rilasciato mostrato al modello

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

File forniti al modello


Caso di studio 7

Genomica regolatoria: misurazione della forza dei loop della cromatina dopo il mascheramento delle varianti strutturali e degli artefatti di mappatura

Quantifica una differenza focale nella forza del loop Hi-C tra casi e controlli dopo aver rimosso dal background dei contatti attesi gli artefatti dovuti alla bassa mappabilità e alle varianti strutturali.

Il loop bersaglio è definito con una risoluzione di 20 kb, ma il modello dei contatti attesi risulta distorto se prima non vengono mascherati i contatti a bassa mappabilità e una striscia SV presente solo nei casi.

Prompt rilasciato mostrato al modello

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

File forniti al modello


Caso di studio 8

Genetica statistica: mappatura QTL multi-parentale con ricostruzione dei fondatori

Mappa un locus per un carattere quantitativo sul cromosoma 1 in una popolazione ricombinante derivata da otto fondatori ricostruendo l'ascendenza dei fondatori prima di testare l’associazione con il fenotipo.

I dati osservabili dei marcatori sono biallelici, ma il segnale biologico è l'ascendenza dei fondatori. Un'analisi difendibile deve quindi ricostruire lo stato dei fondatori, verificare l'orientamento dei marcatori e separare il QTL da un picco spurio allineato al batch.

Prompt rilasciato mostrato al modello

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata

  • founders.tsv.gz: founder alleles at each marker

  • ril_genotypes.npz: observed RIL genotypes (biallelic)

  • phenotypes.tsv.gz: phenotype and covariates

File forniti al modello


Caso di studio 9

Genetica delle popolazioni: ascendenza specifica per genitore e tempistica della mescolanza genetica recente

Inferisci le proporzioni di ascendenza specifiche per ciascun genitore e la tempistica della mescolanza genetica recente a partire da tratti di ascendenza locale fasati, dopo aver corretto gli artefatti reciproci e un'inversione delle etichette specifica di un cromosoma.

Le frazioni di ascendenza e i tempi di admixture cambiano entrambi se gli artefatti dei tratti reciproci, l’inversione delle etichette locali al cromosoma o i denominatori della mappa vengono gestiti in modo errato.

Prompt rilasciato mostrato al modello

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

File forniti al modello


Caso di studio 10

Genetica delle popolazioni: stima della selezione da serie temporali rumorose di DNA antico

Determina quale dei due loci aploidi sia sottoposto a una selezione positiva più intensa a partire da serie temporali antiche delle frequenze alleliche, tenendo conto dell’orientamento allelico, dell’errore direzionale, della deriva genetica e delle variazioni della dimensione della popolazione.

Le traiettorie antiche rumorose non sono direttamente confrontabili finché entrambi i loci non vengono posti sulla stessa scala dell’allele derivato e i valori di errore di sequenziamento a livello di campione forniti non vengono modellati direttamente.

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

File forniti al modello