Pāriet uz galveno saturu
OpenAI

2026. gada 30. jūnijs

Genebench-Pro saturs

Tuvāks ieskats etalonā, tā jautājumos un atbalsta materiālos.

Gadījumu izpētes

Šīs 10 gadījumu izpētes parāda reprezentatīvus GeneBench-Pro jautājumus. Katrā gadījuma izpētē ir iekļauta sākotnējā uzvedne, datu kopas un papildu materiāli. Lai iegūtu pārskatu par etalontestu un galvenajiem secinājumiem, skati paziņojuma bloga ierakstu.

Piezīme: failu priekšskatījumos tiek rādīti pilno datu kopu fragmenti.


Gadījuma izpēte 1

Somatiskā onkoloģija: strukturālo variantu vadīts lēmums par audzēja terapijas ieguvuma-riska attiecību

Novērtē, vai sintētiskam pret TXR1 vērstam inhibitoram ir pozitīva klīniskā lietderība audzējos, kuros mērķa aktivāciju izraisa strukturāls variants. TXR1, TXR1i (TXR1 inhibitors), DLR1 un zvaigznīšu alēļu marķējumi ir sintētiski atsauces marķējumi. 

Mērķa apakšgrupa ir jānosaka, izmantojot garo nolasījumu, ekspresijas, audzēja kvalitātes un farmakogenomikas pierādījumus, pirms ieguvumu un toksicitāti var interpretēt ārstēšanas lēmuma kontekstā.

Publicētā uzvedne, kas parādīta modelim

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Modelim sniegtie faili


Gadījuma izpēte 2

Funkcionālā genomika: CRISPR mērķa validācija: lncRNS (ilgstoši nekodējošā RNS) transkripts vai genomiskais lokuss?

Nosaki, vai šķietama lncRNS atkarība ir transkripta specifiska vai to nosaka tuvējā lokusa un kaimiņgēnu ietekme.

Transkripta virzītiem pierādījumiem jāiztur kontroles pārbaudes attiecībā uz lokālu DNS lokusa traucējumiem, kaimiņgēnu represiju, gidu apmaiņu, GC toksicitāti un plāksnes efektiem.

Publicētā uzvedne, kas parādīta modelim

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.

  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.

  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;

  • more negative numbers indicate stronger loss of fitness;

  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Modelim sniegtie faili


Gadījuma izpēte 3

Statistiskā ģenētika: proteīnu zāļu mērķu prioritizēšana saistītajā ģenētiskajā lokusā

Novērtē tiešo slimības ietekmi uz diviem tuvumā esošiem proteīniem, izmantojot cis daudzmainīgo Mendeļa randomizāciju (cis-MVMR), vienlaikus kontrolējot analīzes skalu, alēļu orientāciju, “uzvarētāja lāstu”, saistības nelīdzsvarotību (LD) un atlikušo lokālo pleiotropiju.

Abiem proteīniem ir kopīgs korelēts lokuss. Analīzei jāpāriet no marginālām asociācijām uz nosacītiem slimības efektiem, kas ņem vērā LD, kopīgā proteīnu mērogā.

Publicētā uzvedne, kas parādīta modelim

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Modelim sniegtie faili


Gadījuma izpēte 4

Klīniskā genomika / nēsātāju skrīnings: DRX1 nēsātāju skrīninga atlikušais risks CNV un pseidogēnu kalibrācijas apstākļos

Novērtē izcelsmes grupu nesēju biežumu, atlikušo risku pēc negatīva skrīninga, partnera nesēju biežumu un risku, ka ieņemtais bērns būs ietekmēts, balstoties uz nesēju skrīninga analīzes datiem.

Atlikušā riska aplēse ir atkarīga no pseidogēnus ņemošiem vērā nēsātāja statusa noteikšanas rezultātiem, dibinātāju haplotipu sapludināšanas, izcelsmei specifiskas analīzes kalibrēšanas un standartizācijas no testētajiem partneriem atpakaļ uz pilno partneru sarakstu.

Publicētā uzvedne, kas parādīta modelim

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

Modelim sniegtie faili


Gadījuma izpēte 5

Single-cell genomics: Activated-monocyte eQTL after ambient RNA correction

Novērtē genotipa ietekmi uz aktivēto monocītu ekspresiju pēc apkārtējās RNS un tehniskās kontaminācijas noņemšanas no vienas šūnas RNA-seq datiem.

Apkārtējā RNS ietekmē gan mērķa ekspresiju, gan marķieru paneli, ko izmanto aktivācijas stāvokļa noteikšanai, tāpēc korekcija jāveic pirms eQTL modeļa.

Publicētā uzvedne, kas parādīta modelim

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

Modelim sniegtie faili


Gadījuma izpēte 6

Strukturālā ģenētika: ligzdotais strukturālais variants: ekspresijas atbalsts un klīniskā saistība

Novērtē, vai anonīmā inversijai līdzīgā lokusā iegultajam strukturālajam subhaplotipam ir kalibrēta klīniskā saistība un ticams ekspresijas atbalsts.

Iegults kopiju devas signāls var radīt neskaidrības plašākas inversijas orientācijas dēļ, tāpēc devas kalibrēšana, ekspresijas atbalsts un klīniskā modelēšana ir jāveic atsevišķi.

Izlaistā uzvedne, ko tu redzi modelis

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Modelim sniegtie faili


Gadījuma izpēte 7

Regulatīvā genomika: hromatīna cilpu stipruma mērīšana pēc strukturālo variantu un kartēšanas artefaktu maskēšanas

Kvantificē fokālas gadījuma–kontroles Hi-C cilpas stipruma atšķirību pēc tam, kad no sagaidāmā kontaktu fona ir noņemti zemas kartējamības un strukturālo variantu artefakti.

Mērķa cilpa ir definēta ar 20 kb izšķirtspēju, taču paredzētais kontaktu modelis tiek izkropļots, ja vispirms netiek maskēti kontakti ar zemu kartējamību un tikai konkrētā gadījumā esošā SV josla.

Publicētā uzvedne, kas parādīta modelim

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Modelim sniegtie faili


Gadījuma izpēte 8

Statistiskā ģenētika: daudzvecāku QTL kartēšana ar dibinātāju rekonstrukciju

Kartē 1. hromosomas kvantitatīvās pazīmes lokusu rekombinantā populācijā ar astoņiem dibinātājiem, rekonstruējot dibinātāju izcelsmi pirms fenotipa asociācijas testēšanas.

Redzamie marķieru dati ir bialēliski, bet bioloģiskais signāls ir dibinātāju senču izcelsme. Tāpēc pamatotai analīzei ir jārekonstruē dibinātāja stāvoklis, jāpārbauda marķiera orientācija un jānošķir QTL no ar partiju sakritīga traucējoša pīķa.

Publicētā uzvedne, kas parādīta modelim

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata

  • founders.tsv.gz: founder alleles at each marker

  • ril_genotypes.npz: observed RIL genotypes (biallelic)

  • phenotypes.tsv.gz: phenotype and covariates

Modelim sniegtie faili


Gadījuma izpēte 9

Populāciju ģenētika: vecāku specifiskā izcelsme un nesenas piejaukšanās laika noteikšana

Izsecini katram vecākam specifiskas senču izcelsmes proporcijas un nesenās ģenētiskās sajaukšanās laiku no fāzētiem lokālās senču izcelsmes posmiem pēc reciproko artefaktu un hromosomai specifiskas marķējumu inversijas izlabošanas.

Senču izcelsmes daļas un impulsu laiki mainās, ja nepareizi apstrādā reciprokos ceļa artefaktus, hromosomu lokālās etiķešu apgriešanas kļūdas vai kartes saucējus.

Publicētā uzvedne, kas parādīta modelim

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Modelim sniegtie faili


Gadījuma izpēte 10

Populāciju ģenētika: atlases novērtēšana no trokšņainām senās DNS laikrindām

Nosaki, kurš no diviem haploīdajiem lokusiem ir pakļauts spēcīgākai pozitīvajai selekcijai, izmantojot senās alēļu frekvenču laikrindas, ņemot vērā alēļu orientāciju, virziena kļūdu, dreifu un mainīgu populācijas lielumu.

Trokšņainas senās trajektorijas nav tieši salīdzināmas, kamēr abi lokusi nav novietoti uz vienas un tās pašas atvasinātās alēles skalas un sniegtās parauga līmeņa sekvencēšanas kļūdu vērtības nav tieši modelētas.

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Modelim sniegtie faili