Přeskoč na hlavní obsah
OpenAI

30. června 2026

Uvnitř Genebench-Pro

Podrobnější pohled na benchmark, jeho otázky a podpůrné materiály.

Případové studie

Těchto 10 případových studií představuje reprezentativní otázky z GeneBench-Pro. Každá případová studie obsahuje původní prompt, datové sady a podpůrné materiály. Přehled benchmarku a hlavní zjištění najdete v oznamovacím blogu.

Poznámka: náhledy souborů zobrazují úryvky z úplných datových sad.


Případová studie 1

Somatická onkologie: rozhodování o poměru přínosů a rizik protinádorové terapie na základě strukturních variant

Odhadni, zda má syntetický inhibitor cílený na TXR1 pozitivní klinický přínos u nádorů, u nichž je aktivace cíle řízena strukturální variantou. TXR1, TXR1i, DLR1 a označení hvězdičkových alel jsou syntetická benchmarková označení. 

Cílovou podskupinu je nutné nejprve odvodit z podkladů ze sekvenování dlouhými čteními, exprese, kvality nádoru a farmakogenomiky, než lze přínos a toxicitu interpretovat jako podklad pro rozhodnutí o léčbě.

Zveřejněný prompt zobrazený modelu

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Soubory poskytnuté modelu


Případová studie 2

Funkční genomika: validace cíle CRISPR: transkript lncRNA nebo genomový lokus?

Rozhodni, zda je zdánlivá závislost na lncRNA specifická pro transkript nebo zda je způsobena vlivem blízkého lokusu a sousedních genů.

Důkazy získané transkripty musí přežít kontroly lokálních poruch lokusu DNA, represe sousedních genů, záměn genů, toxicity plynové chromatografie (GC) a účinků plotny.

Zveřejněný prompt zobrazený modelu

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.

  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.

  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;

  • more negative numbers indicate stronger loss of fitness;

  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Soubory poskytnuté modelu


Případová studie 3

Statistická genetika: prioritizace proteinových cílů léčiv ve vázaném genetickém lokusu

Odhadni přímé účinky dvou blízkých proteinů na onemocnění pomocí cis multivariační Mendelovy randomizace (cis-MVMR) se zohledněním měřítka testu, orientace alel, prokletí vítěze, vazebné nerovnováhy (LD) a reziduální lokální pleiotropie.

Tyto dva proteiny sdílejí korelovaný lokus. Analýza se musí posunout od marginálních asociací k podmíněným účinkům onemocnění, které zohledňují LD, na společné proteinové škále.

Zveřejněný prompt zobrazený modelu

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Soubory poskytnuté modelu


Případová studie 4

Klinická genomika / screening přenašečství: reziduální riziko screeningu přenašečství DRX1 při kalibraci CNV a pseudogenů

Odhadni frekvence přenašečství podle populačního původu, zbytkové riziko přenašečství po negativním výsledku screeningu, pravděpodobnost přenašečství u partnera a riziko, že počatý zárodek/plod bude postižený, a to na základě dat z testu přenašečství.

Odhad reziduálního rizika závisí na vyhodnocení přenašečství zohledňujícím pseudogeny, sloučení zakladatelských haplotypů, kalibraci testu specifické pro ancestrální původ a standardizaci z testovaných partnerů zpět na úplný seznam partnerů.

Zveřejněný prompt zobrazený modelu

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

Soubory poskytnuté modelu


Případová studie 5

Jednobuněčná genomika: eQTL aktivovaných monocytů po korekci ambientní RNA

Odhadni vliv genotypu na expresi aktivovaných monocytů po odstranění okolní RNA a technické kontaminace z dat sekvenování RNA jednotlivých buněk.

Ambientní RNA ovlivňuje jak expresi cíle, tak panel markerů používaný k určení stavu aktivace, takže korekce musí proběhnout před eQTL modelem.

Zveřejněný prompt zobrazený modelu

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

Soubory poskytnuté modelu


Případová studie 6

Strukturní genetika: vnořená strukturní varianta: podpora exprese a klinická asociace

Odhadni, zda vnořený strukturní subhaplotyp uvnitř anonymního inverzního lokusu má kalibrovanou klinickou asociaci a věrohodnou podporu exprese.

Vnořený signál dávky kopií může být zkreslen širší orientací inverze, proto musí kalibrace dávky, expresní podpora a klinické modelování zůstat vzájemně oddělené.

Zveřejněný prompt zobrazený modelu

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Soubory poskytnuté modelu


Případová studie 7

Regulační genomika: měření síly chromatinových smyček po maskování strukturních variant a mapovacích artefaktů

Kvantifikuj rozdíl v síle smyčky Hi-C u fokálních případů a kontrol po odstranění artefaktů s nízkou mapovatelností a strukturálními variantami z očekávaného kontaktního pozadí.

Cílová smyčka je definována v rozlišení 20 kb, ale model očekávaných kontaktů je zkreslený, pokud se nejprve nezamaskují kontakty s nízkou mapovatelností a pruh SV specifický pouze pro daný případ.

Zveřejněný prompt zobrazený modelu

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Soubory poskytnuté modelu


Případová studie 8

Statistická genetika: mapování QTL ve vícerodičovských populacích s rekonstrukcí zakladatelů

Zmapuj lokus kvantitativního znaku chromozomu 1 v rekombinantní populaci osmi zakladatelů rekonstrukcí jejich původu před testováním asociace fenotypu.

Viditelná markerová data jsou biallelická, ale biologickým signálem je zakladatelský původ. Obhajitelná analýza proto musí rekonstruovat zakladatelský stav, ověřit orientaci markerů a oddělit QTL od rušivého píku vázaného na šarži.

Zveřejněný prompt zobrazený modelu

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata

  • founders.tsv.gz: founder alleles at each marker

  • ril_genotypes.npz: observed RIL genotypes (biallelic)

  • phenotypes.tsv.gz: phenotype and covariates

Soubory poskytnuté modelu


Případová studie 9

Populační genetika: rodičovský původ a načasování nedávné příměsi

Odvoď podíly ancestrálního původu specifické pro jednotlivé rodiče a načasování nedávného míšení populací z fázovaných úseků lokálního ancestrálního původu po opravě recipročních artefaktů a inverze štítků specifických pro daný chromozom.

Odhady původových podílů i doby jednotlivých pulzů populačního míšení se mohou změnit, pokud se chybně zohlední artefakty v reciprokých úsecích, lokální prohození ancestrálních označení na chromozomu nebo nesprávné jmenovatele použité v genetické mapě.

Zveřejněný prompt zobrazený modelu

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Soubory poskytnuté modelu


Případová studie 10

Populační genetika: odhad selekce z časových řad starověké DNA se šumem

Odvoď, který ze dvou haploidních lokusů je vystaven silnější pozitivní selekci z dávných časových řad alelových frekvencí, přičemž zohledni orientaci alel, směrovou chybu, genetický drift a měnící se velikost populace.

Starověké trajektorie s vysokým šumem lze porovnávat až po sjednocení obou lokusů na škálu odvozené alely a po přímém zahrnutí vzorkově specifických sekvenačních chyb do modelu.

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Soubory poskytnuté modelu