Overslaan naar hoofdinhoud
OpenAI

30 juni 2026

Binnen Genebench-Pro

Een nadere beschouwing van de benchmark, de bijbehorende vragen en de ondersteunende materialen.

Casestudy's

Deze 10 casestudy’s tonen representatieve vragen uit GeneBench-Pro. Elke casestudy bevat de oorspronkelijke prompt, datasets, en ondersteunend materiaal. Zie de aankondigingsblog voor een overzicht van de benchmark en de belangrijkste bevindingen.

Opmerking: bestandsvoorvertoningen tonen fragmenten uit de volledige datasets.


Case study 1

Somatische oncologie: baten-risicobeslissing voor tumortherapie op basis van structurele varianten

Beoordeel of een synthetische, op TXR1 gerichte remmer positieve klinische waarde heeft bij tumoren waarvan de targetactivatie wordt aangedreven door een structurele variant. TXR1-, TXR1i-, DLR1- en ster-allel-labels zijn synthetische benchmark-labels. 

De doel-subgroep moet worden herleid uit long-read-, expressie-, tumorkwaliteits- en farmacogenomische data voordat de effectiviteit en toxiciteit kunnen worden geïnterpreteerd als een behandelbeslissing.

Vrijgegeven prompt die aan het model wordt getoond

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Bestanden die aan het model zijn verstrekt


Casestudy 2

Functionele genomics: CRISPR-targetvalidatie: lncRNA-transcript of genomische locus?

Beslis of een schijnbare lncRNA-afhankelijkheid transcript-specifiek is of dat deze gedreven wordt door effecten van nabijgelegen loci en naburige genen.

Op transcript-gestuurd bewijs gebaseerde data moeten standhouden tegen controles voor lokale DNA-locusperturbatie, repressie van naburige genen, guideswaps, GC-toxiciteit en plate-effecten.

Vrijgegeven prompt die aan het model wordt getoond

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.

  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.

  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;

  • more negative numbers indicate stronger loss of fitness;

  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Bestanden die aan het model zijn verstrekt


casestudy 3

Statistische genetica: prioriteren van eiwitdoelen voor geneesmiddelen in een gekoppelde genetische locus

Schat de directe ziekte-effecten voor twee nabijgelegen eiwitten met behulp van cis-multivariabele Mendeliaanse randomisatie (cis-MVMR), terwijl er rekening wordt gehouden met de assayschaal, allel-oriëntatie, de 'winner's curse', LD en residuele lokale pleiotropie.

De twee eiwitten delen een geassocieerde locus. De analyse moet verschuiven van marginale associaties naar conditionele, LD-bewuste ziekte-effecten op een gemeenschappelijke eiwitschaal.

Vrijgegeven prompt die aan het model wordt getoond

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Bestanden die aan het model zijn verstrekt


Casestudy 4

Klinische genetica / dragerschapsscreening: Het restrisico bij DRX1-dragerschapsscreening onder kalibratie van CNV's en pseudogenen

Schat herkomstspecifieke dragersfrequenties, het restrisico na een negatieve screening, de dragersfrequentie van de partner en het risico op een aangedane vrucht op basis van assaydata uit dragerscreening.

De schatting van het restrisico hangt af van pseudogeenbewuste dragersbepalingen, founder-haplotype-collaps, herkomstspecifieke assaykalibratie en de standaardisatie van geteste partners terug naar de volledige partnerlijst.

Vrijgegeven prompt die aan het model wordt getoond

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

Bestanden die aan het model zijn verstrekt


Casestudy 5

Single-cell-genomica: eQTL van geactiveerde monocyten na correctie voor omgevings-RNA

Schat een genotype-effect op de expressie van geactiveerde monocyten na verwijdering van omgevings-RNA en technische contaminatie uit single-cell RNA-seq-gegevens.

Ambient RNA beïnvloedt zowel de expressie van het doelwit als het markerpaneel dat wordt gebruikt om de activeringsstatus te bepalen, dus moet correctie plaatsvinden vóór het eQTL-model.

Vrijgegeven prompt die aan het model wordt getoond

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

Bestanden die aan het model zijn verstrekt


Casestudy 6

Structurele genetica: geneste structurele variant: expressie-ondersteuning en klinische associatie

Schat in of een genest structureel subhaplotype binnen een anoniem inversieachtig locus een gekalibreerde klinische associatie en geloofwaardige expressieondersteuning (expression support) heeft.

Een genest kopie-doseringssignaal kan worden verstoord door de bredere inversie-oriëntatie, waardoor dosering-kalibratie, expressie-ondersteuning en klinische modellering strikt gescheiden moeten blijven.

Vrijgegeven prompt die aan het model wordt getoond

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Bestanden die aan het model zijn verstrekt


Casestudy 7

Regulatoire genomica: het meten van chromatineloop-sterkte na maskering van structurele varianten en mappingartefacten

Kwantificeer een focaal case-control-Hi-C-loopsterkteverschil na het verwijderen van artefacten door lage mappability en structurele varianten uit de verwachte contactachtergrond.

De target-loop is gedefinieerd op een resolutie van 20 kb, maar het verwachte contactmodel raakt vervormd tenzij contacten met een lage mappability en een case-specifieke SV-stripe eerst worden gemaskeerd

Vrijgegeven prompt die aan het model wordt getoond

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Bestanden die aan het model zijn verstrekt


Casestudy 8

Statistische genetica: QTL-mapping met meerdere ouders met founderreconstructie

Breng een 'quantitative-trait locus' (QTL) op chromosoom 1 in kaart in een recombinante populatie met acht founders door eerst de founder-afstamming te reconstrueren alvorens de associatie met het fenotype te testen.

De zichtbare markerdata zijn biallelisch, maar het biologische signaal is de founder-afstamming. Een verdedigbare analyse moet daarom de founder-toestand reconstrueren, de markeroriëntatie controleren en de QTL scheiden van een met de batch uitgelijnde verstorende piek (nuisance peak).

Vrijgegeven prompt die aan het model wordt getoond

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata

  • founders.tsv.gz: founder alleles at each marker

  • ril_genotypes.npz: observed RIL genotypes (biallelic)

  • phenotypes.tsv.gz: phenotype and covariates

Bestanden die aan het model zijn verstrekt


Casestudy 9

Populatiegenetica: ouderspecifieke afstamming en de timing van recente vermenging

Leid de ouderspecifieke afstammingsproporties en de timing van recente vermenging af uit gefaseerde lokale afstammingsgebieden, na het herstellen van wederkerige artefacten en een chromosoomspecifieke label-inversie.

Zowel de afstammingsfracties als de vermengingsperioden (pulse times) veranderen als er onjuist wordt omgegaan met wederkerige segment-artefacten, chromosoom-lokale labelinversies of kaartnoemers.

Vrijgegeven prompt die aan het model wordt getoond

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Bestanden die aan het model zijn verstrekt


Casestudy 10

Populatiegenetica: Het schatten van selectie op basis van ruizige tijdreeksen van oud DNA

Leid uit tijdreeksen van prehistorische allel-frequenties af welke van twee haploïde loci onder sterkere positieve selectie staat, terwijl er rekening wordt gehouden met allel-oriëntatie, richtingfouten, genetische drift en een veranderende populatiegrootte.

Ruizige prehistorische trajecten zijn niet direct vergelijkbaar totdat beide loci op dezelfde schaal van afgeleide allelen zijn geplaatst en er direct rekening wordt gehouden met de meegeleverde sequentiefoutwaarden op steekproefniveau.

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Bestanden die aan het model zijn verstrekt