Zum Hauptinhalt springen
OpenAI

30. Juni 2026

Innerhalb von Genebench-Pro

Ein genauerer Blick auf den Benchmark, seine Fragen und die Begleitmaterialien.

Fallstudien

Diese 10 Fallstudien zeigen repräsentative Fragen aus GeneBench-Pro. Jede Fallstudie enthält den ursprünglichen Prompt, Datensätze und unterstützende Materialien. Eine Übersicht über den Benchmark und die wichtigsten Ergebnisse findest du im Ankündigungsblog.

Hinweis: Dateivorschauen zeigen Auszüge aus den vollständigen Datensätzen.


Fallstudie 1

Somatische Onkologie: Nutzen-Risiko-Entscheidung für Tumortherapie anhand struktureller Varianten

Schätze ein, ob ein synthetischer TXR1-gerichteter Inhibitor bei Tumoren, deren Zielaktivierung durch eine strukturelle Variante bedingt ist, einen positiven klinischen Nutzen hat. TXR1, TXR1i, DLR1 und Stern-Allel-Labels sind synthetische Benchmark-Labels. 

Die Zieluntergruppe muss zunächst auf Basis von Long-Read-, Expressions-, Tumorqualitäts- und pharmakogenomischer Evidenz ermittelt werden, bevor Nutzen und Toxizität als Behandlungsentscheidung interpretiert werden können.

Veröffentlichter Prompt, der dem Modell angezeigt wird

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Dem Modell bereitgestellte Dateien


Fallstudie 2

Funktionelle Genomik: CRISPR-Target-Validierung: lncRNA-Transkript oder genomischer Locus?

Entscheide, ob eine scheinbare lncRNA-Abhängigkeit transkriptspezifisch oder durch Effekte nahegelegener Loci und benachbarter Gene bedingt ist.

Transkriptgerichtete Evidenz muss Kontrollen für lokale DNA-Locus-Perturbation, Repression benachbarter Gene, Guide-Swaps, GC-Toxizität und Platteneffekte standhalten.

Veröffentlichter Prompt, der dem Modell angezeigt wird

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.

  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.

  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;

  • more negative numbers indicate stronger loss of fitness;

  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Dem Modell bereitgestellte Dateien


Fallstudie 3

Statistische Genetik: Priorisierung von Protein-Wirkstoffzielen in einem verknüpften genetischen Locus

Schätze direkte Krankheitseffekte für zwei nahe beieinanderliegende Proteine mithilfe der cis-multivariablen Mendelschen Randomisierung (cis-MVMR) unter Berücksichtigung von Assay-Skala, Allelorientierung, Winner’s Curse, LD und residualer lokaler Pleiotropie.

Die beiden Proteine teilen einen korrelierten Locus. Die Analyse muss von marginalen Assoziationen zu bedingten, LD-berücksichtigenden Krankheitseffekten auf einer einheitlichen Proteinskala übergehen.

Veröffentlichter Prompt, der dem Modell angezeigt wird

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Dem Modell bereitgestellte Dateien


Fallstudie 4.

Klinische Genomik / Carrier-Screening: Restrisiko im DRX1-Carrier-Screening unter CNV- und Pseudogen-Kalibrierung

Schätze abstammungsspezifische Trägerfrequenzen, das Restrisiko nach negativem Screening, die Partner-Trägerfrequenz und das Risiko für einen betroffenen Konzeptus anhand von Daten aus Carrier-Screening-Assays.

Die Schätzung des Restrisikos hängt von pseudogenbewussten Trägerbestimmungen, der Zusammenführung von Founder-Haplotypen, einer abstammungsspezifischen Assay-Kalibrierung und der Rückstandardisierung von getesteten Partnern auf die vollständige Partnerliste ab.

Veröffentlichter Prompt, der dem Modell angezeigt wird

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

Dem Modell bereitgestellte Dateien


Fallstudie 5

Einzelzellgenomik: Aktivierter Monozyt-eQTL nach Hintergrund-RNA-Korrektur

Schätze einen Genotyp-Effekt auf die Expression aktivierter Monozyten nach Entfernung von Umgebungs-RNA und technischer Kontamination aus Einzelzell-RNA-seq-Daten.

Umgebungs-RNA beeinflusst sowohl die Zielexpression als auch das Markerpanel, das zur Bestimmung des Aktivierungszustands verwendet wird, daher muss die Korrektur vor dem eQTL-Modell erfolgen.

Veröffentlichter Prompt, der dem Modell angezeigt wird

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

Dem Modell bereitgestellte Dateien


Fallstudie 6

Strukturelle Genetik: Verschachtelte strukturelle Variante: Expressionsunterstützung und klinische Assoziation

Schätze ein, ob ein verschachtelter struktureller Subhaplotyp innerhalb eines anonymen inversionsähnlichen Locus eine kalibrierte klinische Assoziation und eine belastbare Unterstützung durch Expressionsdaten aufweist.

Ein verschachteltes Kopien-Dosierungssignal kann durch die übergeordnete Inversionsorientierung konfundiert werden, daher müssen Dosierungskalibrierung, Expressionsunterstützung und klinische Modellierung voneinander getrennt bleiben.

Veröffentlichter Prompt, der dem Modell angezeigt wird

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Dem Modell bereitgestellte Dateien


Fallstudie 7

Regulatorische Genomik: Messung der Chromatin-Loop-Stärke nach Maskierung von Strukturvarianten und Mapping-Artefakten

Quantifiziere eine fokale Fall-Kontroll-Differenz in der Hi-C-Loop-Stärke, nachdem Artefakte durch geringe Mappability und Strukturvarianten aus dem Hintergrund erwarteter Kontakte entfernt wurden.

Der Ziel-Loop ist mit einer Auflösung von 20 kb definiert, das Modell der erwarteten Kontakte wird jedoch verzerrt, wenn Kontakte mit geringer Mappability und eine fallspezifische SV-Stripe nicht zuerst maskiert werden.

Veröffentlichter Prompt, der dem Modell angezeigt wird

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Dem Modell bereitgestellte Dateien


Fallstudie 8

Statistische Genetik: Multi-Parent-QTL-Mapping mit Founder-Rekonstruktion

Mappe einen Quantitative-Trait-Locus auf Chromosom 1 in einer rekombinanten Population mit acht Founder-Linien, indem du vor dem Testen der Phänotyp-Assoziation die Founder-Abstammung rekonstruierst.

Die sichtbaren Markerdaten sind biallelisch, das biologische Signal ist jedoch die Founder-Abstammung. Eine belastbare Analyse muss daher den Founder-State rekonstruieren, die Marker-Orientierung prüfen und den QTL von einem mit dem Batch korrelierten Störpeak abgrenzen.

Veröffentlichter Prompt, der dem Modell angezeigt wird

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata

  • founders.tsv.gz: founder alleles at each marker

  • ril_genotypes.npz: observed RIL genotypes (biallelic)

  • phenotypes.tsv.gz: phenotype and covariates

Dem Modell bereitgestellte Dateien


Fallstudie 9

Populationsgenetik: Elternteilspezifische Abstammung und Datierung rezenter Admixture

Schätze elternteilspezifische Abstammungsanteile und die Datierung rezenter Admixture aus phasierten lokalen Abstammungssegmenten nach Korrektur reziproker Artefakte und einer chromosomenspezifischen Label-Inversion.

Abstammungsanteile und Pulszeiten ändern sich beide, wenn reziproke Segmentartefakte, chromosomenlokale Label-Inversionen oder Kartennenner falsch behandelt werden.

Veröffentlichter Prompt, der dem Modell angezeigt wird

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Dem Modell bereitgestellte Dateien


Fallstudie 10

Populationsgenetik: Schätzung von Selektion aus verrauschten Zeitreihen alter DNA

Leite aus alten Allelfrequenz-Zeitreihen ab, welcher von zwei haploiden Loci stärker positiver Selektion unterliegt, und berücksichtige dabei Allel-Orientierung, Richtungsfehler, Drift und veränderte Populationsgröße.

Verrauschte alte Trajektorien sind erst direkt vergleichbar, wenn beide Loci auf dieselbe Skala abgeleiteter Allele gebracht und die bereitgestellten Sequenzierungsfehlerwerte auf Probenebene direkt modelliert werden.

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Dem Modell bereitgestellte Dateien