مرکزی مواد پر جائیں
OpenAI

۳۰ جون، ۲۰۲۶

Genebench-Pro کے اندر

بینچ مارک، اس کے سوالات اور معاون مواد کا تفصیلی جائزہ.

کیس اسٹڈیز

یہ دس کیس اسٹڈیز GeneBench-Pro کے نمائندہ سوالات پیش کرتی ہیں. ہر کیس اسٹڈی میں اصل پرومپٹ، ڈیٹا سیٹس اور معاون مواد شامل ہوتے ہیں. بینچ مارک اور اہم نتائج کے جائزے کے لیے، اعلان بلاگ دیکھیں.

نوٹ: فائل کے پیش نظارے مکمل ڈیٹا سیٹس سے اقتباسات دکھاتے ہیں.


کیس اسٹڈی 1

سومیٹک آنکولوجی: ساختی ویرینٹ کی رہنمائی میں ٹیومر تھراپی کے فائدہ-خطرہ کا فیصلہ

یہ اندازہ لگائیں کہ آیا TXR1 کو نشانہ بنانے والا ایک مصنوعی انہیبیٹر اُن ٹیومرز میں مثبت طبی افادیت رکھتا ہے جن میں ہدف کی فعّالیت ایک ساختی ویریئنٹ کے باعث ہوتی ہے. TXR1، TXR1i، DLR1 اور سٹار ایلیل لیبل مصنوعی بینچ مارک لیبل ہیں. 

علاج کے فیصلے کے طور پر فائدے اور زہریت کی تشریح سے پہلے ہدفی ذیلی گروہ کو لانگ-ریڈ، اظہاری، ٹیومر معیار اور فارماکوجینومک شواہد سے حاصل کرنا ہوگا.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

A molecular tumor board registry contains trial-eligible advanced solid-tumor cases considered for a TXR1-directed inhibitor. Estimate, for tumors with SV-driven TXR1 target-mediated activation at time zero, the marginal effect of TXR1i versus non-TXR1 systemic therapy on week-16 clinical benefit as if all patients had an assessable week-16 visit. Also estimate the 8-week treatment-limiting toxicity/discontinuation risk under TXR1i in the same target population. Report net clinical utility = benefit risk difference (percentage points) - 0.35 * toxicity risk (percentage points), and choose therapy_class_code 1 if TXR1i has positive net utility and 0 otherwise. 

Use percentage-point units for all non-code quantities. Positive benefit means TXR1i improves week-16 clinical benefit relative to non-TXR1 systemic therapy.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"therapy_class_code": <int>,
4
"benefit_rd_pp": <float>,
5
"toxicity_dropout_risk_pp": <float>,
6
"net_clinical_utility_pp": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 2

فنکشنل جینومکس: CRISPR ہدف کی توثیق: lncRNA ٹرانسکرپٹ یا جینیاتی مقام؟

یہ طے کریں کہ آیا lncRNA پر بظاہر انحصار ٹرانسکرپٹ-مخصوص ہے یا قریبی لوکس اور پڑوسی جین کے اثرات کی وجہ سے پیدا ہوا ہے.

ٹرانسکرپٹ-ہدایت یافتہ شواہد کو مقامی DNA-لوکس میں خلل، پڑوسی جین کی سرکوبی، گائیڈ کی تبدیلی، GC زہریلے پن اور پلیٹ اثرات کے کنٹرولز کے بعد بھی برقرار رہنا ہوگا.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

You are given pooled CRISPRi screening data, guide-level local expression measurements, transcript-targeting CasRx follow-up data, and single-guide follow-up growth measurements for a nominated lncRNA program (LINC473) and a nearby coding gene (KIN1). The identifiers LINC473, KIN1, and ANKRD42 are synthetic benchmark labels; any resemblance to real human genes is coincidental.

Estimate the requested quantities.

Definitions:

  • lncrna_specific_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% effective knockdown of the dominant LINC473 transcript, not local DNA-locus effects.

  • neighbor_mediated_lfc: the pooled-screen matched-control day-10 log2 growth effect expected at 100% KIN1 repression in the local LINC473-locus model after accounting for concomitant LINC473 transcript knockdown.

  • advance_target: 1 if the evidence supports advancing LINC473 as a transcript-directed target, else 0.

Conventions:

  • all growth effects are log2(day10/day0) competitive-growth effects relative to matched controls;

  • more negative numbers indicate stronger loss of fitness;

  • set advance_target to 1 only if lncrna_specific_lfc <= -0.08 and neighbor_mediated_lfc > -0.25; otherwise 0.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"advance_target": <int>,
4
"lncrna_specific_lfc": <float>,
5
"neighbor_mediated_lfc": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 3

شماریاتی جینیات: ایک مربوط جینیاتی مقام میں پروٹین دوا کے اہداف کو ترجیح دینا

سیس ملٹی ویری ایبل مینڈیلین رینڈمائزیشن (cis-MVMR) کا استعمال کرتے ہوئے دو قریبی پروٹینوں کے لیے براہ راست بیماری کے اثرات کا تخمینہ لگائیں جب کہ پرکھ کے پیمانے، ایلیل اورینٹیشن، ونر کرس، LD، اور بقایا مقامی پیلیوٹروپی کو ہینڈل کریں.

دونوں پروٹینز ایک باہم مربوط لوکس کا اشتراک کرتے ہیں. تجزیے کو حاشیائی وابستگیوں سے آگے بڑھ کر مشترکہ پروٹین پیمانے پر مشروط، LD کو مدنظر رکھنے والے بیماری کے اثرات تک منتقل ہونا ہوگا.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

You are given association summary statistics and metadata for two nearby proteins (PROTA and PROTB), a binary disease outcome, a locus correlation reference, and protein measurement records.

Goal: estimate the direct log-odds effect of each protein on the disease outcome per +1 SD increase in log10 concentration, conditional on the other protein.

Interpretation: theta_PROTA and theta_PROTB use the same log-odds per-SD scale defined in the goal.

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"theta_PROTA": <float>,
4
"theta_PROTB": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 4

کلینیکل جینومکس / کیریئر اسکریننگ: CNV اور سیوڈوجین کیلیبریشن کے تحت DRX1 کیریئر اسکریننگ کا باقی ماندہ خطرہ

حامل اسکریننگ جانچ کے ڈیٹا سے نسب-مخصوص حامل تعددات، منفی اسکریننگ کے بعد باقی ماندہ خطرہ، شریکِ حیات میں حامل ہونے کی تعدد اور متاثرہ حاصلِ حمل کا خطرہ تخمینہ کریں.

بقایا خطرے کا تخمینہ سیوڈوجین کو مدنظر رکھنے والی کیریئر کالز، فاؤنڈر-ہیپلوٹائپ کے ادغام، نسب-مخصوص اسّے کی کیلیبریشن اور ٹیسٹ کیے گئے پارٹنرز سے مکمل پارٹنر فہرست تک کی جانے والی معیار بندی پر منحصر ہوتا ہے.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

Using cohort_roster.tsv.gz, partner_roster.tsv.gz, calibration_controls.tsv.gz, target_metadata.tsv.gz, and assay_observations.tsv.gz, estimate residual reproductive risk for an autosomal recessive DRX1 condition. Report all quantities on the probability scale, not as percentages: carrier_frequency_afr and carrier_frequency_eur among screening-roster adults; residual_carrier_risk_afr_negative for an AFR screening-roster adult with a negative DRX1 screen; partner_carrier_frequency_full_roster for a uniformly sampled partner_roster.tsv.gz row; and couple_reproductive_risk for an affected conceptus when the index person is AFR and screen-negative and the partner is drawn from partner_roster.tsv.gz. Assume autosomal recessive inheritance with a 1/4 affected-conceptus risk conditional on both biological parents being carriers. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"carrier_frequency_afr": <float>,
4
"carrier_frequency_eur": <float>,
5
"residual_carrier_risk_afr_negative": <float>,
6
"partner_carrier_frequency_full_roster": <float>,
7
"couple_reproductive_risk": <float>
8
},
9
"reasoning": "<description of method and QC>"
10
}

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 5

سنگل سیل جینومکس: محیطی RNA کی تصحیح کے بعد فعال شدہ مونو سائٹ eQTL

سنگل سیل RNA-سیکوئنس ڈیٹا سے محیطی RNA اور تکنیکی آلودگی کو ہٹانے کے بعد فعال شدہ مونو سائٹ کے اظہار پر جینوٹائپ کے اثر کا تخمینہ لگائیں.

ایمبیئنٹ RNA ہدف کے اظہار اور ایکٹیویشن کی حالت کا تعین کرنے کے لیے استعمال ہونے والے مارکر پینل دونوں کو متاثر کرتا ہے، اس لیے تصحیح eQTL ماڈل سے پہلے ہونی چاہیے.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

Estimate the per-allele log rate ratio for CXCL10 expression in the activated monocyte subpopulation from the provided single-cell RNA-seq data. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"beta_activated": <float>
4
},
5
"reasoning": "<description of method and QC>"
6
}

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 6

ساختی جینیات: متداخل ساختی تغیر: اظہار کی معاونت اور طبی تعلق

اندازہ لگائیں کہ آیا کسی غیر شناخت شدہ انورژن نما لوکس کے اندر موجود تو در تو ساختی ذیلی ہیپلوٹائپ کی معیّر شدہ طبی وابستگی اور قابلِ اعتماد اظہار کی تائید موجود ہے.

ایک تہ در تہ کاپی-ڈوزیج سگنل وسیع تر اِنورژن کی سمت بندی سے گڈمڈ ہو سکتا ہے، اس لیے ڈوزیج کیلیبریشن، ایکسپریشن سپورٹ اور کلینیکل ماڈلنگ کو الگ الگ رکھنا ضروری ہے.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

Analyze the released files for anonymous Locus Q. Estimate the full-cohort source-population clinical association and molecular expression support for the calibrated nested segment-B structural copy dosage, separating the nested segment-B dosage from the broader outer-orientation dosage. Report subhap_log_or as the natural-log source-population total-effect odds ratio for case status per additional calibrated segment-B copy. Report expression_log_fc as the natural-log expression fold-change per calibrated segment-B copy for the expression-supported gene. Report target_support_code as 1 if the supported gene has a positive expression_log_fc and the clinical association is protective (subhap_log_or < 0), otherwise 0. Report n_calibrated_carriers as the number of reliable breakpoint-panel samples carrying at least one segment-B copy. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"n_calibrated_carriers": <int>,
4
"target_support_code": <int>,
5
"expression_log_fc": <float>,
6
"subhap_log_or": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 7

ریگولیٹری جینومکس: ساختی ویریئنٹ اور میپنگ آرٹیفیکٹ کی ماسکنگ کے بعد کرومیٹن لوپ کی مضبوطی کی پیمائش

متوقع رابطے کے پس منظر سے کم میپ ایبلٹی اور ساختی ویرینٹ آرٹیفیکٹس کو ہٹانے کے بعد ایک فوکل کیس-کنٹرول Hi-C لوپ اسٹرینتھ کے فرق کی مقدار متعین کریں.

ٹارگٹ لوپ 20 kb ریزولوشن پر متعین ہے، لیکن متوقع کانٹیکٹ ماڈل مسخ ہو جاتا ہے جب تک کہ کم میپیبلٹی والے کانٹیکٹس اور صرف کیس والی SV اسٹرائپ کو پہلے ماسک نہ کیا جائے.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

You are given Hi-C contact matrices at 20 kb and 40 kb resolution plus bin annotations. Estimate the loop enrichment at the 20 kb interaction between `bin_id = 8` and `bin_id = 17` in `bins_20kb.tsv.gz`. Report three quantities: `case_loop_strength` (mean log2(observed/expected) across case replicates), `control_loop_strength` (mean log2(observed/expected) across control replicates), and `delta_loop_strength` (case minus control).

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"case_loop_strength": <float>,
4
"control_loop_strength": <float>,
5
"delta_loop_strength": <float>
6
},
7
"reasoning": "<description of method and QC>"
8
}

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 8

شماریاتی جینیات: بانی کی تعمیر نو کے ساتھ کثیر والدین QTL نقشہ سازی

فینوٹائپ کی وابستگی کی جانچ سے پہلے بانی نسب کی بازتعمیر کر کے، آٹھ بانیوں والی بازترکیبی آبادی میں کروموسوم-1 کے مقداری صفت لوکس کا نقشہ بنائیں.

مرئی مارکر ڈیٹا دو-ایلیلی ہے، لیکن حیاتیاتی سگنل بانی نسب ہے. لہٰذا ایک قابلِ دفاع تجزیہ کو بانی حالت کی بازسازی کرنی، مارکر کی سمت کی جانچ کرنی اور QTL کو بیچ سے ہم آہنگ غیر مطلوبہ چوٹی سے الگ کرنا ہوگا.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

Map the chromosome 1 QTL in an 8-founder multi-parent population. Report the position (cM) and which founder carries the high-effect allele.

Report high_founder as "F1".."F8".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"high_founder": "<string>",
4
"qtl_pos_cM": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

Data files:

  • markers.tsv.gz: marker metadata

  • founders.tsv.gz: founder alleles at each marker

  • ril_genotypes.npz: observed RIL genotypes (biallelic)

  • phenotypes.tsv.gz: phenotype and covariates

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 9

آبادیاتی جینیات: والدین کی مخصوص آبائی نسل اور حالیہ آمیزش کے وقت کا تعین

باہمی آرٹیفیکٹس اور کروموسوم-مخصوص لیبل کی الٹ پھیر کو درست کرنے کے بعد، فیز شدہ مقامی نسبی قطعات سے والدین-مخصوص نسبی تناسبات اور حالیہ اختلاطِ نسب کے وقت کا استنباط کریں.

اگر متقابل ٹریکٹ آرٹیفیکٹس، کروموسوم-مقامی لیبل کی الٹ پھیر، یا نقشہ کے مخرجوں کو غلط طریقے سے سنبھالا جائے تو نسبی تناسبات اور پلس اوقات دونوں تبدیل ہو جاتے ہیں.

ماڈل کو دکھائی گئی جاری کردہ پرومپٹ

You are given phased local-ancestry tracts for one admixed individual. Estimate, for each transmitted parental haplotype, the fraction of ancestry A across the called tract span and the number of generations since a single recent admixture pulse. Label parent1 as the haplotype with the smaller ancestry-A fraction and parent2 as the haplotype with the larger ancestry-A fraction. 

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"parent1_A_fraction": <float>,
4
"parent1_t": <float>,
5
"parent2_A_fraction": <float>,
6
"parent2_t": <float>
7
},
8
"reasoning": "<description of method and QC>"
9
}

ماڈل کو فراہم کی گئی فائلیں


کیس اسٹڈی 10

آبادیاتی جینیات: شور زدہ قدیم ڈی این اے کے وقتی سلسلوں سے انتخاب کا تخمینہ

قدیم ایلیل-فریکوئنسی زمانی سلسلوں سے یہ استنباط کریں کہ دو ہیپلائیڈ لوکائی میں سے کون سا زیادہ مضبوط مثبت انتخاب کے زیرِ اثر ہے، جبکہ ایلیل کی سمت بندی، سمتی خرابی، ڈرفٹ اور بدلتے ہوئے آبادی کے حجم کو ملحوظ رکھا جائے.

شور زدہ قدیم راستے براہِ راست قابلِ موازنہ نہیں ہوتے جب تک کہ دونوں لوکس کو ایک ہی مشتق-ایلیل پیمانے پر نہ رکھا جائے اور فراہم کردہ نمونہ سطح کی سیکوینسنگ-ایرر کی قدروں کو براہِ راست ماڈل نہ کیا جائے.

You are given allele-frequency time series data from two haploid loci sampled over multiple generations.

One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.

Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.

The selected_locus value must be "A" or "B".

These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.

Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:

JSON

1
{
2
"answer": {
3
"selected_locus": "<string>",
4
"s": <float>
5
},
6
"reasoning": "<description of method and QC>"
7
}

ماڈل کو فراہم کی گئی فائلیں