Pasar al contenido principal
OpenAI

30 de junio de 2026

Dentro de Genebench-Pro

Una mirada más cercana al referente, sus preguntas y materiales de apoyo.

Casos de estudio

Estos 10 casos de estudio presentan preguntas representativas de GeneBench-Pro. Cada caso de estudio incluye el prompt original, los conjuntos de datos y los materiales de apoyo. Para obtener una descripción general del punto de referencia y los hallazgos clave, consulta el blog de anuncio.

Nota: Las vistas previas de los archivos muestran fragmentos de los conjuntos de datos completos.


Caso de estudio 1

Oncología somática: decisión de beneficio-riesgo de la terapia antitumoral guiada por variantes estructurales

Estima si un inhibidor sintético dirigido a TXR1 tiene utilidad clínica positiva en tumores cuya activación de la diana está impulsada por una variante estructural. TXR1, TXR1i, DLR1 y las etiquetas star-allele son etiquetas de referencia sintéticas. 

El subgrupo objetivo debe recuperarse a partir de evidencia de lecturas largas, de expresión, de calidad tumoral y farmacogenómica antes de que el beneficio y la toxicidad puedan interpretarse como base para una decisión de tratamiento.

Prompt publicado que se muestra al modelo

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Archivos proporcionados al modelo

patient_idanalysis_setagesexsitecalendar_periodecogtumor_burdenprior_linesprior_resistancelineage_classtherapy_classassessed16benefit16tox_stop_8wktime_zero_day
MTB0001173.8MS1P220.78731ATXR1i010
MTB0002155.2MS3P112.63701ATXR1i1000
MTB0003168.8FS4P200.89121ATXR1i1110
MTB0004182.8FS2P224.10100BTXR1i1000
MTB0005165.5FS1P317.011ATXR1i1000

Covariables del registro, terapia, evaluación de la semana 16, beneficio y toxicidad temprana.


Caso de estudio 2

Genómica funcional: validación de dianas CRISPR: ¿transcrito de lncRNA o locus genómico?

Determina si una aparente dependencia de lncRNA es específica del transcrito o si está impulsada por efectos del locus cercano y de genes vecinos.

La evidencia dirigida por transcritos debe resistir los controles de perturbación local del locus de ADN, represión de genes vecinos, intercambios de guías, toxicidad por GC y efectos de placa.

Prompt publicado que se muestra al modelo

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.
  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.
  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;
  • more negative numbers indicate stronger loss of fitness;
  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Archivos proporcionados al modelo

guide_idnominal_targetchrcoordstranddist_lnc_tss_bpdist_neighbor_tss_bpguide_gc_frac
g001LINC473chr7100014+14300.624
g002LINC473chr7100035-43670.584
g003LINC473chr7100051+116560.622
g004LINC473chr7100066-59660.617
g005LINC473chr7100088+74770.715

Guía, coordenadas, objetivos, distancias y funciones de GC.


Caso de estudio 3

Genética estadística: priorización de blancos farmacológicos proteicos en un locus genético vinculado

Estima los efectos directos sobre la enfermedad de dos proteínas cercanas mediante aleatorización mendeliana multivariable cis (cis-MVMR), teniendo en cuenta la escala del ensayo, la orientación de los alelos, el sesgo del ganador, el desequilibrio de ligamiento (LD) y la pleiotropía local residual.

Las dos proteínas comparten un locus correlacionado. El análisis debe pasar de asociaciones marginales a efectos de la enfermedad condicionales, considerando el LD, en una escala proteica común.

Prompt publicado que se muestra al modelo

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Archivos proporcionados al modelo

snppos_bpeffect_alleleother_allelemafbetasepval
rs20000050000000AC0.422150.0064386683107068080.0032673300912034120.04876727714241972
rs20000150010126AC0.057090.0110089933375813010.0069552392087504070.11345916603941006
rs20000250020253GT0.090210.0099220147571163190.0056330230270155180.07817048492026045
rs20000350030379GT0.483990.0105692156141645730.00322914197402374450.0010638520681901973
rs20000450040506AG0.377030.0070365513782386540.00332975923212698020.034580976884336506

Resúmenes de asociaciones de proteínas en etapa de tamizaje para PROTA.


Caso de estudio 4

Genómica clínica / tamizaje de portadores: riesgo residual del tamizaje de portadores de DRX1 con calibración de CNV y pseudogenes

Estima las frecuencias de portadores específicas por ascendencia, el riesgo residual tras un resultado negativo en el cribado, la frecuencia de portadores de la pareja y el riesgo de que el producto de la concepción esté afectado a partir de datos de un ensayo de cribado de portadores

La estimación del riesgo residual depende de las llamadas de portadores conscientes de pseudogenes, el colapso de haplotipos fundadores, la calibración del ensayo específica por ascendencia y la estandarización desde los socios evaluados hacia la lista completa de socios.

Prompt publicado que se muestra al modelo

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

Archivos proporcionados al modelo

sample_idcolecciónascendenciafamily_history_tier
S_EUR_0001screeningEUR0
S_EUR_0002screeningEUR0
S_EUR_0003screeningEUR0
S_EUR_0004screeningEUR0
S_EUR_0005screeningEUR1

Adultos en el registro de tamizaje con ascendencia y contexto del tamizaje.


Caso de estudio 5

Genómica de célula única: eQTL de monocitos activados tras la corrección de ARN ambiental

Estimar un efecto del genotipo sobre la expresión en monocitos activados después de eliminar el ARN ambiental y la contaminación técnica de datos de RNA-seq de célula única.

El ARN ambiental afecta tanto la expresión del objetivo como el panel de marcadores usado para determinar el estado de activación, por lo que la corrección debe realizarse antes del modelo de eQTL.

Prompt publicado que se muestra al modelo

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

Archivos proporcionados al modelo

cell_iddonortotal_umiHBBIFI6ISG15LST1CXCL10
D01_C001D011113734835
D01_C002D01110363311210
D01_C003D0111419812639
D01_C004D01125076043217
D01_C005D0110459125115

Conteos de UMI por célula para genes marcadores, marcadores de contaminación y el gen objetivo.


Caso de estudio 6

Genética estructural: variante estructural anidada: evidencia de expresión y asociación clínica

Estimar si un subhaplotipo estructural anidado dentro de un locus anónimo similar a una inversión tiene una asociación clínica calibrada y evidencia de expresión confiable.

Una señal anidada de dosificación de copias puede confundirse con la orientación más amplia de la inversión, por lo que la calibración de la dosificación, el soporte de expresión y el modelado clínico deben mantenerse separados.

Prompt publicado que se muestra al modelo

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Archivos proporcionados al modelo

sample_idcaseageage_bandsexpc1pc2pc3ancestry_groupclinic_stratumrecruitment_stream
Q00012150.4550_640-1.01514-0.21032-0.08849EURtertiaryclinic
Q00028057.3950_640-1.25987-0.124980.2344EURregionalregistry
Q00029168.465_plus00.915980.621770.01891AFRtertiaryclinic
Q00030174.0765_plus10.21125-0.59634-0.08197EAScommunityregistry
Q00032182.8265_plus0-1.12034-0.243720.14665EURcommunityclinic

Datos clínicos y de covariables de la cohorte completa.


Caso de estudio 7

Genómica regulatoria: medición de la fuerza de los bucles de cromatina tras el enmascaramiento de variantes estructurales y artefactos de mapeo

Determina la diferencia en la intensidad de un bucle de Hi-C específico entre casos y controles tras eliminar del fondo de contactos esperados los artefactos debidos a la baja capacidad de mapeo y a las variantes estructurales.

El bucle objetivo está definido con una resolución de 20 kb, pero el modelo de contactos esperados se distorsiona si antes no se enmascaran los contactos en regiones difíciles de mapear y una franja asociada a una variante estructural presente solo en los casos.

Prompt publicado que se muestra al modelo

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

Archivos proporcionados al modelo

bin_idchromstartendgc_contentmappabilityre_sites
0chr84000004200000.461990338215725940.97875742147042735
1chr84200004400000.50441242085346770.89010849434983975
2chr84400004600000.432184515849381940.90568792893267123
3chr84600004800000.47331972826812180.93765298406647893
4chr84800005000000.44449560621507480.86825655179818774

Anotaciones de segmentos (bins) a la resolución objetivo


Caso de estudio 8

Genética estadística: mapeo de QTL multiparental con reconstrucción de fundadores

Identifica un locus de rasgo cuantitativo en el cromosoma 1 en una población recombinante de ocho fundadores, reconstruyendo su ascendencia antes de evaluar la asociación con el fenotipo.

Los datos de marcadores observables son bialélicos, pero la señal biológica es la ascendencia de los fundadores. Por lo tanto, un análisis riguroso debe reconstruir el estado de los fundadores, verificar la orientación de los marcadores y distinguir el QTL de un pico de ruido asociado al lote.

Prompt publicado que se muestra al modelo

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata
  • founders.tsv.gz: founder alleles at each marker
  • ril_genotypes.npz: observed RIL genotypes (biallelic)
  • phenotypes.tsv.gz: phenotype and covariates

Archivos proporcionados al modelo

marker_idchrpos_cM
m2_065259.762431265596575
m2_103294.52656615104739
m2_107298.18761427503033
m2_079272.20130244108847
m1_054149.907510212292195

Identificadores de marcadores, cromosomas y posiciones en el mapa genético.


Caso de estudio 9

Genética de poblaciones: ascendencia por progenitor y estimación del momento de una mezcla reciente

Inferir las proporciones de ascendencia específicas de cada progenitor y el momento de la mezcla reciente a partir de segmentos con fase asignada de ascendencia local, después de corregir artefactos recíprocos y una inversión de etiquetas específica para un cromosoma.

Las fracciones de ascendencia y los tiempos de pulso cambian si los artefactos de tramos recíprocos, la inversión local de etiquetas por cromosoma o los denominadores del mapa se manejan de forma incorrecta.

Prompt publicado que se muestra al modelo

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

Archivos proporcionados al modelo

chromhapstart_morganend_morganancposteriorlow_complexity_frac
chr1h10.030.505A0.9850.08
chr1h10.5050.535B0.620.92
chr1h10.5351.478849A0.9850.08
chr1h11.5037271.852681B0.9850.08
chr1h11.8526812.422373A0.9850.08

Tramos de ascendencia local con fases y coordenadas, etiquetas de ascendencia, valores posteriores y anotaciones de control de calidad.


Caso de estudio 10

Genética de poblaciones: estimación de la selección a partir de series temporales ruidosas de ADN antiguo.

Inferir cuál de dos loci haploides está bajo una selección positiva más fuerte a partir de series temporales antiguas de frecuencia alélica, considerando la orientación de los alelos, el error direccional, la deriva genética y los cambios en el tamaño de la población.

Las trayectorias antiguas con ruido no son directamente comparables hasta que ambos loci se colocan en la misma escala de alelos derivados y se modelan directamente los valores de error de secuenciación a nivel de muestra proporcionados.

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Archivos proporcionados al modelo

generationalt_readstotal_readsseq_errorsample_year
636400.16-4500
1234450.16-4 278
1841550.16-4 056
2438700.16-3 833
3036900.16-3 611

Serie temporal de recuento de lecturas para el locus A.