Прескокни до главната содржина
OpenAI

30 јуни 2026 г.

Во Genebench-Pro

Подетален поглед на реперот, неговите прашања и придружните материјали.

Студии на случај

Овие 10 студии на случај прикажуваат репрезентативни прашања од GeneBench-Pro. Секоја студија на случај ги вклучува оригиналниот промпт, збирките на податоци и придружните материјали. За преглед на реперот и клучните наоди, погледнете ја најавениот блог.

Забелешка: Прегледите на датотеките прикажуваат извадоци од целосните збирки податоци.


Студија на случај 1

Соматска онкологија: одлука за односот корист-ризик при терапија на тумор водена од структурни варијанти

Проценете дали синтетички инхибитор насочен кон TXR1 има позитивна клиничка корист кај тумори кај кои активацијата на целта е поттикната од структурна варијанта. TXR1, TXR1i, DLR1 и ознаките star-allele се синтетички референтни ознаки. 

Целната подгрупа мора да се реконструира врз основа на докази од долги прочитувања, експресија, квалитет на туморот и фармакогеномски докази пред користа и токсичноста да можат да се протолкуваат како одлука за третман.

Објавен промпт прикажан на моделот

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Фајлови доставени за модел

patient_idanalysis_setagesexsitecalendar_periodecogtumor_burdenprior_linesprior_resistancelineage_classtherapy_classassessed16benefit16tox_stop_8wktime_zero_day
MTB00011738MS1P22078731ATXR1i010
MTB00021552MS3P11263701ATXR1i1000
MTB00031688FS4P20089121ATXR1i1110
MTB00041828FS2P22410100BTXR1i1000
MTB00051655FS1P317011ATXR1i1000

Регистарски коваријати, терапија, проценка во 16-та недела, придобивка и рана токсичност.


Студија на случај 2

Функционална геномика: валидација на таргетот со CRISPR: lncRNA транскрипт или геномски локус?

Одредете дали привидната зависност од lncRNA е специфична за транскриптот или е условена од ефекти на блиски локуси и соседни гени.

Доказите насочени од транскриптот мора да останат валидни по контролите за локална пертурбација на ДНК-локусот, репресија на соседни гени, замени на водичи, GC-токсичност и ефекти од плочата.

Објавениот промпт прикажан на моделот

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.
  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.
  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;
  • more negative numbers indicate stronger loss of fitness;
  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Датотеки пратени за модел

guide_idnominal_targetchrcoordstranddist_lnc_tss_bpdist_neighbor_tss_bpguide_gc_frac
g001LINC473chr7100014+14300.624
g002LINC473chr7100035-43670.584
g003LINC473chr7100051+116560.622
g004LINC473chr7100066-59660.617
g005LINC473chr7100088+74770.715

Водич за координати, цели, растојанија и функции на GC.


Студија на случај 3

Статистичка генетика: Приоритизирање на протеински терапевтски цели во поврзан генетски локус

Проценете ги директните ефекти на болеста за два блиски протеини со користење цис мултиваријабилна Менделова рандомизација (cis-MVMR), притоа справувајќи се со скалата на анализата, ориентацијата на алелите, проклетството на победникот, нерамнотежата на врзаност (LD) и резидуалната локална плејотропија.

Двата протеина споделуваат корелиран локус. Анализата треба да премине од маргинални асоцијации кон условни ефекти на болеста што ја земаат предвид LD, на заедничка протеинска скала.

Објавениот промпт прикажан на моделот

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Датотеки доставени за модел

snppos_bpeffect_alleleother_allelemafbetasepval
rs20000050000000AC0.422150.0064386683107068080.0032673300912034120.04876727714241972
rs20000150010126AC0.057090.0110089933375813010.0069552392087504070.11345916603941006
rs20000250020253GT0.090210.0099220147571163190.0056330230270155180.07817048492026045
rs20000350030379GT0.483990.0105692156141645730.00322914197402374450.0010638520681901973
rs20000450040506AG0.377030.0070365513782386540.00332975923212698020.034580976884336506

Резиме на протеински асоцијации во фазата на скрининг за PROTA.


Студија на случај 4

Клиничка геномика / скрининг за носителство: преостанат ризик при скрининг за носителство на DRX1 со калибрација за CNV и псевдогени

Проценете ги фреквенциите на носителство специфични за потеклото, резидуалниот ризик по негативен скрининг, фреквенцијата на носителство кај партнерот и ризикот за засегнат концептус врз основа на податоци од анализа за скрининг на носителство.

Проценката на резидуалниот ризик зависи од определувањата на носителство што ги земаат предвид псевдогените, обединувањето на основачките хаплотипови, калибрацијата на анализата специфична за потеклото и стандардизацијата од тестираните партнери назад кон целосниот список на партнери.

Објавениот промпт што му е прикажан на моделот

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

Датотеки доставени за модел

sample_idcollectionancestryfamily_history_tier
S_EUR_0001screeningEUR0
S_EUR_0002screeningEUR0
S_EUR_0003screeningEUR0
S_EUR_0004screeningEUR0
S_EUR_0005screeningEUR1

Возрасни лица од списокот за скрининг со потекло и контекст за скрининг.


Студија на случај 5

Едноклеточна геномика: eQTL за активирани моноцити по корекција на амбиентална РНК

Проценете го ефектот на генотипот врз експресијата кај активираните моноцити по отстранување на амбиенталната РНК и техничката контаминација од податоците од едноклеточно RNA-seq.

Амбиенталната РНК влијае и врз експресијата на целта и врз панелот маркери што се користи за одредување на состојбата на активација, па корекцијата мора да се изврши пред да се примени eQTL моделот.

Објавен промпт прикажан на моделот

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

Датотеки доставени на модел

cell_iddonortotal_umiHBBIFI6ISG15LST1CXCL10
D01_C001D011113734835
D01_C002D01110363311210
D01_C003D0111419812639
D01_C004D01125076043217
D01_C005D0110459125115

Броења на UMI по клетка за маркерските гени, маркерите за контаминација и целниот ген.


Студија на случај 6

Структурна генетика – вгнездена структурна варијанта: поддршка на експресијата и клиничка поврзаност

Проценете дали вгнезден структурен субхаплотип во анонимен локус сличен на инверзија има калибрирана клиничка поврзаност и веродостојна експресиска поддршка.

Вгнезден сигнал за дозирање на копии може да биде конфундиран од пошироката ориентација на инверзијата, па затоа калибрацијата на дозирањето, експресиската поддршка и клиничкото моделирање мора да останат одделни.

Објавен промпт прикажан на моделот

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Датотеки пратени за модел

ид_на_примерокслучајвозраствозрасна_групаполpc1pc2pc3група_на_потеклоклинички_слојканал_за_регрутација
Q000121504550_640-101514-021032-008849EURтерцијаренклиника
Q000280573950_640-125987-01249802344EURрегионаленрегистар
Q00029168465_plus0091598062177001891AFRтерцијаренклиника
Q000301740765_plus1021125-059634-008197EASзаедницарегистар
Q000321828265_plus0-112034-024372014665EURзаедницаклиника

Клинички и коваријатни податоци за целата кохорта.


Студија на случај 7

Регулаторна геномика: Мерење на јачината на хроматинските јамки по маскирање на структурни варијанти и артефакти од мапирање

Квантифицирајте ја фокалната разлика во јачината на Hi-C јамката помеѓу случајот и контролата по отстранување на артефактите од ниска мапабилност и структурни варијанти од позадината на очекувани контакти.

Целната јамка е дефинирана со резолуција од 20 Kb, но моделот на очекувани контакти е изобличен ако прво не се маскираат контактите со ниска мапабилност и лентата SV што е присутна само во случајот.

Објавен промпт прикажан на моделот

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Датотеки испратени за модел

bin_idchromstartendgc_contentmappabilityre_sites
0chr84000004200000.461990338215725940.97875742147042735
1chr84200004400000.50441242085346770.89010849434983975
2chr84400004600000.432184515849381940.90568792893267123
3chr84600004800000.47331972826812180.93765298406647893
4chr84800005000000.44449560621507480.86825655179818774

Анотации за бинови со целна резолуција.


Студија на случај 8

Статистичка генетика: QTL-мапирање со повеќе родители и реконструкција на основачите

Мапирајте локус за квантитативна особина на хромозом 1 во рекомбинантна популација со осум основачи, со реконструирање на потеклото од основачите пред да ја тестирате фенотипската асоцијација.

Видливите податоци од маркерите се биалелни, но биолошкиот сигнал е потекло од основачите. Затоа, аргументирана анализа мора да ја реконструира основачката состојба, да ја провери ориентацијата на маркерите и да го одвои QTL од непожелен пик усогласен со серијата податоци.

Објавен промпт прикажан на моделот

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata
  • founders.tsv.gz: founder alleles at each marker
  • ril_genotypes.npz: observed RIL genotypes (biallelic)
  • phenotypes.tsv.gz: phenotype and covariates

Датотеки пратени на модел

marker_idchrpos_cM
m2_065259.762431265596575
m2_103294.52656615104739
m2_107298.18761427503033
m2_079272.20130244108847
m1_054149.907510212292195

Идентификатори на маркери, хромозоми и позиции на генетската карта.


Студија на случај 9

Популациска генетика: потекло специфично за родители и време на неодамнешно мешање

Проценете ги пропорциите на потекло специфично за родителите и времето на неодамнешно генетско мешање од фазирани сегменти на локално потекло, по коригирање на реципрочните артефакти и инверзија на ознаките специфична за хромозомата.

Делови од потеклото и времињата на пулсот се менуваат ако артефактите од реципрочните траки, инверзијата на локалните ознаки на хромозомот или именителите на мапата се обработуваат неправилно.

Објавениот промпт прикажан на моделот

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Фајлови доставени за модел

chromhapstart_morganend_morganancposteriorlow_complexity_frac
chr1h10.030.505A0.9850.08
chr1h10.5050.535B0.620.92
chr1h10.5351.478849A0.9850.08
chr1h11.5037271.852681B0.9850.08
chr1h11.8526812.422373A0.9850.08

Фазни сегменти на локално потекло со координати, ознаки за потекло, апостериорни вредности и забелешки за контрола на квалитет.


Студија на случај 10

Популациона генетика: Проценување на селекцијата од зашумени временски серии на древна ДНК

Изведете заклучок кој од двата хаплоидни локуси е под посилна позитивна селекција врз основа на древни временски серии на фреквенции на алели, земајќи ги предвид ориентацијата на алелите, насочената грешка, генетскиот дрифт и променливата големина на популацијата.

Шумните древни траектории не се директно споредливи сè додека двата локуса не се постават на иста скала на изведен алел и додека обезбедените вредности за грешка при секвенционирање на ниво на примерок не се моделираат директно.

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Датотеки пратени за модел

generationalt_readstotal_readsseq_errorsample_year
636400.16-4500
12.34.45016-4278
18.41.55016-4056
24.38.70016-3833
30.36.90016-3611

Временска серија на бројот на читања за локусот A.