Independent development audit · exon-level CNV/LGA

LGASieve cannot yet support a clinical call.

A source-preserving reconstruction, method correction, nested cross-validation, 903,168-trial stochastic simulation, information-limit analysis, public truth-set audit, and literature comparison. The strongest answer is a defensible NO-GO for autonomous clinical or scientific release—and a concrete route to a valid screening assay.

Report date: 2026-07-31Assay context: 50-target BRCA1/BRCA2 count sliceFinal verification: 128 passed, 3 skipped, 1 warning, 7 subtests passed in 48.74sLockbox: sealed
NO-GODo not issue autonomous reportable BRCA CNV calls from this pipeline or this 50-target normalization domain.

The conclusion does not depend on one bad threshold. At the predeclared family alpha of 0.05, the best new development model reached 61.6% deterministic synthetic sensitivity, but only 9.9% for single-exon duplications. With count sampling added, single-exon duplication sensitivity was 10.2%; 50%-mosaic duplication sensitivity was 0.9%. The 147 native libraries are presumed—not verified—negatives, so their clean fraction is not clinical specificity.

903,168stochastic trials; each test host held out from training and calibration
49.6%single-exon germline DEL sensitivity, stochastic
10.2%single-exon germline DUP sensitivity, stochastic
96.9%presumed-negative clean fraction; explicitly not specificity
“Perfect” performance is not a reasonable stopping rule. Repeated tuning on one development collection can make internal numbers look perfect while worsening transport. The valid stopping rule is: freeze a method that passes predeclared development gates, then test it once on independent, orthogonally confirmed specimens. No method tested here earned that lockbox opening.
ClaimEvidence hereAllowed wordingNot allowed
Sampling powerKnown-baseline Poisson oracleOptimistic information ceilingClinical sensitivity
Algorithm behaviorHeld-out deterministic and stochastic count perturbationsInternal software sensitivityEnd-to-end assay sensitivity or LoD
False-flag burden147 previously examined 1000 Genomes librariesPresumed-negative clean fractionClinical specificity
External benchmarkPublished callers on different full capture panelsArchitecture prior and plausibility checkTransferable performance for this assay
Reproducible workflow

What I did, in order

  1. Preserved and reconstructed the source. Seven Dropbox archives were retained with SHA-256 hashes. The legacy reference path was restored as a byte-identical copy of reference_grch37; no algorithm was changed for baseline reproduction.
  2. Reproduced the original implementation. The untouched baseline completed with 107 passed, 3 expected skips, and 7 subtests. The archived decision was already NO_GO_CLINICAL_OR_SCIENTIFIC_RELEASE.
  3. Located and corrected one proposal-loss defect. The window scan formerly rejected a real multi-exon interval if any internal exon had the opposite residual sign, and then silently retained only the strongest 50 candidates. The corrected scan tests every contiguous window whose two boundaries support the direction; a regression test covers a three-exon deletion with one noisy internal exon.
  4. Re-ran the frozen native partition. The wrapper can now load and validate the archived 50/60/37 partition instead of reconstructing an unavailable selection file. The caller completed all 97 non-reference libraries.
  5. Built a candidate-excluded conditional Gaussian experiment. Every one of 147 hosts was outer-held-out; all preprocessing and robust conformal family calibration were fitted inside training folds. All 462 callable windows were scored.
  6. Recursively swept operating points. Family alpha 0.01 through 0.40, joint two-sided max-statistic calibration, and an exploratory direction-weighted allocation were compared. No high/high operating point appeared.
  7. Added stochastic LGA mimics. Binomial thinning modeled deletions; coupled Poisson increments modeled duplications. Germline and 50%-mosaic amplitudes, five widths, whole-gene-feasible events, eight common-random-number replicates, depth strata, exact and reciprocal-overlap localization, and specimen-cluster uncertainty were tested.
  8. Separated physical information from implementation loss. Exact Poisson ceilings, signal geometry, empirical residual floors, target correlation, reference learning, batch/population proxy transport, and calibration resolution were audited.
  9. Benchmarked against literature and public truth resources. The 2025 12-caller benchmark supplement was preserved locally; ICR96, panelcnDataset, GIAB, simulator packages, professional standards, and MLPA confirmation guidance were evaluated for applicability.
  10. Kept the lockbox sealed. Since development gates failed, the 75-specimen holdout was not downloaded or analyzed.
Preserved archiveBytesSHA-256
exomedepth_development_v2.zip401,811527eb7cebfa811fc02cc1cf09ddf1742eedcd7b40e2f9fbdfe3d235986b35852
lgasieve.zip1,226,46798c3caacd3a1eaf5b95d13829379e8b2f2623073cd24a2b898ab5e2b69f4b554
reference_grch37.zip214,2742665e9fb40e1f3ffc98ed851665639f6dd08cfe241adb05e62bc600d87aec773
scripts.zip363,439d7629a71e455c8adf097d406880ef22a49b17cb695023c22182814e27c2b23cb
tests.zip821,210080a7d228473626b3498f9ac662b9f993b2b570493c571260a1c0e20dd1f9c32
validation_corrected.zip3,667,430653a8aec481580ba4de8e229cc2339476c607599309825e67a30ef0bd14cd66a
validation_v2.zip422,371711e32c7937ec947c1fedd358bade94f9a03b8a24a4781e341a007d2a3df28f7

Dropbox source folder: authenticated folder. Archive hashes establish the reconstruction boundary; the working tree contains the documented research changes below.

Root-cause analysis

The dominant failure is information recovery, not nominal depth alone

1. Count sampling is often adequate

Under an optimistic known-baseline Poisson oracle with family alpha split over 46 targets and two directions, mean single-exon power is 93.5% for a heterozygous deletion and 87.5% for a single-copy duplication. In libraries with median depth at least 100, oracle duplication power is 98.9%.

This is an optimistic sampling ceiling, not clinical or modeled sensitivity. It proves only that low observed duplication sensitivity cannot be blamed entirely on shot noise.

2. The model discards most duplication information

The older nested predictive model converted only 6.3% of the high-depth oracle duplication power into detections. The new conditional model improves multi-exon behavior, but still detects only 10.2% of stochastic single-exon germline duplications at alpha 0.05.

This gap points to residual capture/batch variation, reference transport, target-specific calibration, and multiple-testing cost—not simply “sequence deeper.”
0.585log2 signal for 3 copies vs −1.000 for a heterozygous deletion
1.70×Poisson Bhattacharyya depth multiplier for equal DUP/DEL information
2.92×equal-z Gaussian depth multiplier; model-dependent upper comparison
46/50callable targets; the archived workflows also contain a conflicting 48-target mask

Measured residual structure

DiagnosticObservedInterpretation
Median target residual MAD0.182 log2Too large for reliable single-copy duplication separation at strict family thresholds.
Median Poisson delta-method SD0.113 log2Empirical residual noise is 1.62× the Poisson component at the median target.
Quadrature systematic floor0.140 log2Extra depth cannot remove this estimated capture/model component.
Count vs target residual MADSpearman ρ = -0.818Depth helps, but does not eliminate the residual floor.
Adjacent target correlationmedian r = 0.199Multi-exon windows contain less independent information than their width suggests.
Leave-super-population-out transport4.7% worseA weak proxy only; actual site/run/kit/lot/operator metadata are absent.
60-calibrator two-direction p-value floor3.28%Resolution and threshold stability are inadequate for a 1% family false-positive claim.
Whole-gene non-identifiability. BRCA2 occupies 27 of 50 targets. A uniform BRCA2 deletion changes the panel median to 0.5; median normalization makes affected BRCA2 targets exactly 1.0 and makes unaffected BRCA1 targets appear doubled. A duplication similarly makes BRCA2 neutral and BRCA1 appear deleted. This is an identifiability failure: no amount of depth or additional references can recover an absolute scale that was normalized away. Whole-gene and panel-global events must be no-called unless external invariant anchors are added.
Mask ambiguity is a release blocker. The source matrix contains 50 targets; this audit finds 46 callable (BRCA2_E01, BRCA1_E23, BRCA1_E11, BRCA1_E01 excluded), while archived validation paths sometimes use 48. One versioned mask must be frozen and propagated through training, calibration, evaluation, reporting, and truth matching.
Recursive in-silico testing

More realistic simulation confirms the failure—and reveals instability

The primary simulation perturbs only the outer-held-out host. Reference selection, covariance estimation, robust score scaling, and the empirical sample-family null are learned without that host. Eight common-random-number replicates make method comparisons less noisy. The specimen is the bootstrap cluster; treating 903,168 correlated trials as independent would give a misleadingly narrow CI.

Stochastic sensitivity by event widthSensitivity improves with event width, while duplications and mosaic events remain substantially harder.0%25%50%75%100%Germline DEL: 49.6%Germline DUP: 10.2%50% mosaic DEL: 3.2%50% mosaic DUP: 0.9%1Germline DEL: 80.7%Germline DUP: 23.0%50% mosaic DEL: 6.0%50% mosaic DUP: 1.7%2Germline DEL: 91.6%Germline DUP: 40.3%50% mosaic DEL: 10.6%50% mosaic DUP: 2.6%3Germline DEL: 98.8%Germline DUP: 71.8%50% mosaic DEL: 28.5%50% mosaic DUP: 8.2%5Germline DEL: 100.0%Germline DUP: 85.6%50% mosaic DEL: 42.1%50% mosaic DUP: 14.1%7Number of affected exonsGermline DELGermline DUP50% mosaic DEL50% mosaic DUP
Stochastic overlap sensitivity at family alpha 0.05. Deletion uses coupled binomial thinning; duplication uses coupled Poisson increments. “50% mosaic” means a 0.75/1.25 expected count factor, not a clinically validated mosaic LoD.
Event widthGermline DELGermline DUP50% mosaic DEL50% mosaic DUP
1 exon49.6%10.2%3.2%0.9%
2 exons80.7%23.0%6.0%1.7%
3 exons91.6%40.3%10.6%2.6%
5 exons98.8%71.8%28.5%8.2%
7 exons100.0%85.6%42.1%14.1%
Uncertainty decomposition. Overall stochastic overlap sensitivity was 35.88%. Its specimen-cluster 95% interval was 34.61%–37.07%, whereas the finite Monte Carlo interval conditional on the fixed grid was only 35.849%–35.905%. Host variation, not simulation count, controls uncertainty.
Transport instability. One outer fold achieved only about 21.6% overlap while the other four were approximately 39.2–39.9%. The same fold was weak in deterministic testing. This reproducible fold effect is a red flag for reference and calibration transport.

Operating-point recursion

Increasing alpha raises sensitivity only by accepting more native flags. Direction-weighted Bonferroni calibration (20% DEL / 80% DUP) did not repair the imbalance: at alpha 0.05 it sacrificed most deletion power and still failed to make single-exon duplication acceptable.

Sensitivity and clean-fraction operating curveAs alpha increases, simulated sensitivity rises but presumed-negative clean fraction falls. No point satisfies both high sensitivity and high clean fraction.0%25%50%75%100%0.010.020.050.100.200.40Family alphaPresumed-negative clean fractionOverall synthetic sensitivitySingle-exon DELSingle-exon DUP
Deterministic conditional-Gaussian development operating curve. Clean fraction uses presumed negatives and is not specificity. Sensitivity uses exact 0.5/1.5 count factors and is not clinical sensitivity.
Family αNative cleanOverall syntheticSingle DELSingle DUP
0.0197.3%15.7%5.7%0.4%
0.0296.6%34.4%19.3%3.4%
0.0595.9%61.6%48.1%9.9%
0.1090.5%78.4%75.9%24.1%
0.2076.9%88.2%91.3%45.5%
0.4055.8%93.6%96.7%65.2%
Simulation coverage boundary. These count-level perturbations test normalization, scoring, localization, calibration, and sampling variability. They do not reproduce hybrid-capture allele dropout, GC/mappability changes, pseudogene misalignment, breakpoint position, duplicate marking, UMI consensus, insert-size effects, library preparation, or variant truth ascertainment. BAM/FASTQ injection with Bamgineer/SECNVs and physical contrived specimens are the next layers, not substitutes for each other.
Code change and fresh rerun

The proposal defect was real, but it was not the main bottleneck

BehaviorBeforeAfterVerification
Internal noisy exonAny wrong-sign internal z-score vetoed the exact multi-exon intervalOnly both boundaries must support the direction; internal noise is permittedNew BRCA2 E12–E14 deletion regression with a wrong-sign E13
Candidate familySorted and silently truncated to 50All boundary-qualified contiguous windows retainedComplete small-panel family; max-statistic remains downstream
Frozen partitionRerun depended on a missing selection fileArchived 50/60/37 JSON accepted after exact membership, size, uniqueness, and overlap validationPositive load test and duplicate-membership rejection test
97/97native non-reference libraries completed
94/97presumed-negative libraries without final flags
35/97libraries with more proposals after the scan fix
3libraries with 4 total final flags; unchanged from the archived rerun

The candidate set increased in 35 libraries (range +1 to +23 candidates), yet the final four flags were unchanged. That A/B result is valuable: proposal loss existed, but normalization, residual modeling, corroboration, and calibration are the limiting layers now.

The observed clean fraction is 96.91% (exact 95% interval 91.23%–99.36%). These samples lack orthogonal negative confirmation and were previously examined; this cannot be reported as specificity.

Iteration ledger

IterationDesignInternal resultDecision
Prior corrected caller516 derivatives, only 8 host librariessensitivity 52.3%; clean fraction 75.0%Rejected: low sensitivity and heuristic score bypass
Beta-binomial, reused PoN null95 PoN libraries reused for LOO null; 4 dev hostssensitivity 55.4%; clean fraction 100.0%Rejected as optimistic reuse bias
Beta-binomial, strict split50 fit / 60 calibration / 8 dev hosts; 4,704 spikessensitivity 7.7%; clean fraction 100.0%Rejected: sensitivity 7.7%
Predictive residualNested 5×5 cross-fit; 147 hosts; 24,108 spikessensitivity 47.6%; clean fraction 96.6%Rejected: single-DUP sensitivity 7.5%
Official ExomeDepth defaultv1.1.16; 4 hosts; 544 spikessensitivity 6.6%; clean fraction 100.0%Rejected: default HMM sensitivity 6.6%
Official ExomeDepth tunedtransition=0.01; development-tuned; 4 native hostssensitivity 43.0%; clean fraction 100.0%Rejected: sensitivity 43.0%; negative n=4
Patched LGASieve native rerun50 PoN fit; 97 non-reference development librariesclean fraction 96.9%Improved false-flag burden; still below ≥99% gate
Conditional OAS / conformal modelNested 5×5; 147 outer-held hosts; 56,448 exact-factor eventsoverall 61.6%; single DUP 9.9%Rejected: no high/high operating point; exploratory only
Stochastic conditional model903,168 trials; 147 hosts; 8 replicatesoverall 35.9%; single DUP 10.2%; mosaic single DUP 0.9%Rejected: poor DUP/mosaic sensitivity and fold transport
Published evidence and truth data

Successful callers use much more normalization context than this slice provides

The strongest external comparison is the 2025 benchmark of 12 germline CNV callers on four real, MLPA-prevalidated targeted-panel datasets (495 samples, 231 CNVs). Its supplementary archive is preserved locally. At aggregate gene level, the best tools performed well—but they normalized using complete capture panels with many more regions of interest. Those numbers are architecture evidence, not performance claims for LGASieve.

CallerTPFPFNAggregate sensitivityWhy it does not transfer
CoNVaDING228150398.70%High sensitivity, lower precision; full panel context
GATK-gCNV22427796.97%Cohort latent model and thousands of intervals
DECoN220621195.24%Optimized reference sets over the captured panel
CODEX2219931294.81%Latent-factor/batch structure estimated from broad coverage
ExomeDepth205382688.74%Designed to build an aggregate reference from many exome/panel bins
A recurring design rule: analyze all usable on-target and off-target regions to establish library scale and latent technical factors, then mask output to the reportable genes. panelcn.MOPS explicitly excludes the gene of interest when selecting correlated references. Restricting normalization itself to BRCA1/2 removes the very anchors needed for gene-wide events.

Public truth-set audit

ResourceTruthAccessUse here
ICR9696 TruSight Cancer v2 samples; 66 positive, 30 negative; 68 validated CNVs (25 single-exon, 43 multi-exon; 51 DEL, 17 DUP); MLPA confirmationFASTQ/BAM data under EGA controlled access EGAS00001002428 / EGAD00001003335Valuable external algorithm benchmark after authorization, but different capture assay and not clinical validation of LGASieve
panelcnDataset161 analyzed samples from a 170-sample TruSight Cancer collection; 41 CNVs (13 single-exon, 28 multi-exon; 36 DEL, 5 DUP)EGA controlled access EGAS00001002481 / EGAD00001003400Useful for method transport research; not assay-matched and not openly downloadable
NIST/GIAB nstd175Genome-wide structural-variant benchmark regions/callsPublic truth; matching raw assay libraries still requiredUseful to test truth matching and some large events, but not an exon-level BRCA capture truth set
Current 1000 Genomes count matrixNo orthogonal CNV-negative confirmation in supplied filesLocalDevelopment residuals and false-flag burden only
ifCNV repository dataTSCA and Juno somatic tumor panels with gene-level aCGH labels and no exon breakpoint truthDownloaded and preserved at commit cf11324af004c31624f0668d77c88a89bcbb4b77Rejected as a BRCA germline exon-level validation set; retained only for possible domain-shift research
No openly accessible, assay-matched, orthogonally confirmed BRCA1/2 positive-and-negative FASTQ cohort was found that can turn this run into clinical validation. Downloading a different panel’s dataset can test software transport; it cannot validate this library preparation, capture design, laboratory process, or reportable range.
Recommended solution

Redesign the measurement and calibration domain, then simplify the caller

LayerRequired changeWhyRelease control
Assay inputsRetain all capture-panel on-target bins, validated off-target bins, explicit invariant control probes, and/or molecule-count spike-insRestores absolute scale and latent technical contextWhole-gene calls are no-call until external anchors are validated
MetadataRecord site, run, flow cell, kit, lot, capture batch, read length, operator, extraction, input mass, and UMI metricsAllows batch-aware reference selection and true transport testsOut-of-domain batches no-call automatically
Reference cohortBuild an assay/batch-matched pool; estimate learning curves inside training data; current data suggest roughly 100 candidates before plateau, not a universal constantSmall reused PoNs make empirical tails coarse and unstableNo patient/replicate derivative may cross train/calibration/test
NormalizationCandidate-excluded hierarchical negative-binomial/beta-binomial latent-factor model with target, library, batch, GC, insert-size, duplicate, and capture effectsPrevents the candidate from defining its own baseline and models overdispersion directlyVersion all masks, factors, and hyperparameters
StatisticJoint DEL/DUP window likelihood over all callable runs; covariance-aware matched filter; no per-exon hard vetoUses coherent multi-exon evidence without silently losing exact boundariesOne frozen family definition and one max-statistic calibration
CalibrationBatch-aware cross-conformal or held-out max-statistic calibration with at least ~300 verified negatives for a perfect-run lower bound near 99%60 calibrators cannot resolve a 1% two-direction family error reliablyReport confidence intervals and no-calls; never tune on the lockbox
Orthogonal evidenceBAF from validated heterozygous sites with site-specific beta-binomial noise; split/discordant reads, local assembly, and UMI evidence as secondary supportIndependent channels can distinguish dosage from capture artifactsAbsence of breakpoint evidence must not veto exon-scale events
ReportingUse NGS as a high-sensitivity screen; reflex every candidate to MLPA or an independently designed dosage assayFalse positives are tolerable only when they cannot become final clinical callsSingle-probe/isolated exon findings require a second kit or different technique
Pragmatic near-term product. A sensitivity-first screening layer with mandatory reflex confirmation is achievable before an autonomous caller. Its validated claim is “selects specimens for orthogonal dosage testing,” not “clinically confirms a BRCA CNV.”

Why this is more likely to work

ExomeDepth, DECoN, panelcn.MOPS, GATK-gCNV, CODEX2, CNVkit, SavvyCNV, ClearCNV, and Cobalt differ statistically, but the successful implementations share broad normalization context, reference/QC selection, explicit latent or count-noise modeling, and validation on physical truth. None supplies evidence that a BRCA-only 50-target median-scaled slice can identify whole-gene dosage autonomously.

Verification plan

A validation ladder that prevents simulation from becoming a clinical claim

  1. Unit and invariant tests. Freeze target ordering/mask, coordinate conventions, candidate exclusion, boundary behavior, joint family size, p-value monotonicity, duplicate handling, deterministic seeds, and atomic artifacts.
  2. Analytic limits. Re-run exact Poisson/negative-binomial power, compositional identifiability, calibration resolution, and reference learning curves for every assay revision.
  3. Count-level stochastic simulation. Use hierarchical negative-binomial libraries with fitted target/batch covariance, germline and mosaic DEL/DUP, heterogeneous exon amplitudes, GC/insert perturbations, missingness, contamination, and out-of-domain batches. Hold the host and all derivatives out.
  4. BAM/FASTQ simulation. Inject CNVs with Bamgineer, SECNVs, or an independently verified simulator before alignment; include pseudogene/homology and breakpoint placement. Compare expected and observed molecule/read-count shifts.
  5. Physical analytic specimens. Use reference materials, cell-line mixtures, engineered controls, and dilution series spanning single exon, multi-exon, whole gene, DEL, DUP, mosaic fraction, input mass, and depth. Include low-quality and interference conditions.
  6. Clinical truth cohort. Orthogonally confirmed positives and verified negatives across the reportable range; MLPA or independent dosage testing establishes truth. Do not label unflagged population samples as negatives.
  7. Repeatability and reproducibility. Replicate operators, instruments, sites, kits, lots, captures, runs, days, and bioinformatics deployments. Cluster intervals by patient and original library, not by synthetic derivative.
  8. Frozen prospective holdout. Lock code, container, model, masks, gates, and truth-matching rules; open once. If it fails, the set becomes development data and a new holdout is required.

Predeclared gates for an autonomous reportable caller

GateMinimumDenominator / interval ruleCurrent status
Analysis completion≥99%All eligible physical specimens; no silent exclusions97/97 dev only
Overall positive agreement≥95%; one-sided patient-cluster lower bound ≥90%Orthogonally confirmed positives, stratified by type/widthNot met internally
Single-exon DEL≥95%Physical truth; sufficient n for interval, not repeated derivatives49.6% stochastic
Single-exon DUP≥90%Physical truth; target- and depth-stratified10.2% stochastic
Multi-exon DEL and DUP≥95% eachWidth, gene, breakpoint, and whole-gene strataNot met in each stratum
Negative agreement≥99%; one-sided lower bound ≥95%Verified negatives; patient/batch clustering; no-call included transparentlyNot assessable
Reproducibility≥95% call and localization agreementLots/sites/operators/runs; predefined equivocal handlingArchived 12.5–25%
No-call ratePredeclared by use caseIncluded in completion and clinical workflow impactMust be designed

How many negatives?

If every tested negative is clean, exact one-sided 95% lower bounds require at least 59 specimens to exceed 95%, 299 to exceed 99%, and 2995 to exceed 99.9%. These are ideal independent-specimen counts; batch clustering and any failure increase the requirement.

The AMP/API/CAP in-silico recommendations support simulation as a supplement and gap-filler, not as a replacement for physical clinical specimens. EMQN, ACMG, and CCMG guidance likewise requires validation across the reportable laboratory workflow and clear limitations.
Holdout governance

The lockbox remains sealed

SEALED_NOT_DOWNLOADED_NOT_ANALYZEDNo lockbox specimen was downloaded, inspected, scored, or used for tuning
75 specimenssame-project technical holdout; not a verified clinical cohort

The canonical specimen-list hash is 37f1939ddff65d89b969a8308cf089d28d517d062ea09b9f9a8d1d499bdc550a. Opening this set now would spend the only holdout on a method that already fails development gates. The correct action is to keep it sealed, build the broader assay/reference model, freeze the resulting pipeline, and open the lockbox once only if all development gates pass. Even a passing result would be technical evidence, not clinical validation, because the truth status is not verified.

Audit trail

Artifacts, hashes, and exact rerun commands

ArtifactBytesSHA-256
SOURCE_MANIFEST.json2,6756b0c886cf24d9e293b617eed8277c96fd41895ec6f991f5225a2ed2b6b4a9e77
lgasieve/segment.py13,70154ec5b510986f9e5fe1a1ddd3a784e6ec261a60d3f61a4d303d1b39c5fdf482c
lgasieve/config.py16,593c552a1e2e0dd25972a2e8aec250d6bdbe8fb6e9a3008b99c9d4d115e84c98bfa
scripts/native_reanalysis.py9,131cc2579b9ea90927ccb470283759e5711efecfbaf94481dfc08adfd30db3a2921
scripts/conditional_gaussian_experiment.py39,652ef3468e1bbd1036c0ad98d469c43a3a0f49429ef714461ca578d6c298fefd77a
scripts/stochastic_conditional_experiment.py38,485d28c2e9c94f3b6296a2451518701285e11e06c6b420877a6faa020114a637e17
scripts/information_limit_audit.py12,582e5dc458dbcbf75ecfdf1114c8b94f7cbfb5983490c9a00495379ae7a3af288b4
scripts/residual_structure_audit.py38,5523a444996013feece91d2330d98c743eeeb6115900136fcbf3c022a50c7b2bbb5
scripts/build_research_report.py78,2916bc391a28a3fc33f3d44a220d67a1cc07e2f70c9037c7e4b60cd4f86ccfbab49
tests/test_science_regressions.py13,622348bf7e53bdc6552e11da4498db06e1a81e7394cc3c2b80a4efe2fe1855401e0
tests/test_conditional_gaussian.py6,8134df355440acce566379f6ce5279a3c067bcb36b60a04c1abda1c1b6800276d53
tests/test_stochastic_conditional.py4,9542a7be41359cc433059be4ca01501a681d66451104c0d5a65c635e17fac5a8fd6
tests/test_information_limit_audit.py1,437c8752dd2b82d1442d5c207973794df472883a7f4543403d5dddc21f6a9de5ad7
tests/test_residual_structure_audit.py6,024c02b0e39a13d01aaf22ef0d5e918e52019466aaffd909db62b79af6e03bf644c
analysis/native_reanalysis_complete_scan.json26,8509d50c0738632143994c03632d2691ed054a70ea25dbdb274cbad5c8ae6c4a369
analysis/conditional_gaussian_nested.json67,606adeb39cee448cfd4abe3e1840b7d0a8af78484c0cd42cf2ea10482ceadf9fceb
analysis/stochastic_conditional_nested.json612,9563d8d4ed7aef2e8724f2a20780988f782dc18a735bc99133be81837ff9333e841
analysis/information_limit_audit.json6,5950e47fc7ccaf458d1b3502dc6b4e8cc79876d6dea6d6b8222683d7e3358e26c93
analysis/residual_structure_audit.json40,973f1043dac1fb0bab13310bc8cd9157821f5ee822468cd87f421ac175cb96327c7
PROJECT_STATUS.md7,75353dac296f54f07c28472f3e7e5fc864c267f9e208be7f3aec005524c074c7cfa
validation_v2/lockbox_protocol.json2,856c275179159c412f0ada761934f21f3fd8c62dc0813010d35dea1e2154550a4df
validation_v2/revalidation_summary.json11,843ba33139b0751ae252745b6197ea701c8abcb61b9e56555bd6456fab80463b63f
external_data/archives/supplementary_data_bbae645.zip2,891,42828daa22ff7f4cac99b05ec37df851f03e5ff7e5dc9454e79e3d10ca1d76f0cc1
Exact principal commands

python -m pytest -q

python -m scripts.native_reanalysis --partition-json work\exomedepth_development_v2\partition.json --output analysis\native_reanalysis_complete_scan.json

python -m scripts.conditional_gaussian_experiment --output analysis\conditional_gaussian_nested.json

python -m scripts.stochastic_conditional_experiment --output analysis\stochastic_conditional_nested.json

python -m scripts.information_limit_audit --output analysis\information_limit_audit.json

python -m scripts.residual_structure_audit --output analysis\residual_structure_audit.json

python -m scripts.build_research_report --test-summary "128 passed, 3 skipped, 1 warning, 7 subtests passed in 48.74s"

Reproducibility boundary. Source archives and key inputs are hashed. JSON artifacts embed source-script and input hashes where implemented. Timestamps and elapsed runtime can vary; scientific summaries, fixed-seed simulations, partitions, and hash-linked source should not.
Primary sources

Literature and standards consulted

All web sources accessed 2026-07-31. Links point to the paper, regulator/standards page, archive record, or manufacturer documentation rather than a search result.

  1. Munté et al. (2025), systematic comparison of 12 germline CNV callers across four real targeted-panel datasets.
  2. Plagnol et al. (2012), ExomeDepth: read-depth CNV calling with an optimized aggregate reference.
  3. Fowler et al. (2016/2017), DECoN development and validation.
  4. Povysil et al. (2017), panelcn.MOPS for targeted panels.
  5. Babadi et al. (2023), GATK-gCNV scalable probabilistic CNV calling in 7,962 exomes.
  6. Talevich et al. (2016), CNVkit and the use of on- and off-target reads.
  7. Roca et al. (2019), Atlas-CNV targeted-panel validation.
  8. Laver et al. (2022), SavvyCNV and off-target coverage.
  9. Jiang et al. (2018), CODEX2 and latent-factor/batch modeling.
  10. Hartmann et al. (2022), ClearCNV.
  11. Bergmann et al. (2022), Cobalt clinical exome CNV caller.
  12. Bamgineer: controlled CNV/allele-specific BAM simulation.
  13. SECNVs: simulation of CNVs in exome-sequencing data.
  14. AMP/API/CAP recommendations for in-silico approaches in clinical NGS validation.
  15. EMQN best-practice guidelines for hereditary breast and ovarian cancer testing.
  16. ACMG technical standard for BRCA1/BRCA2 testing.
  17. CCMG laboratory guideline: end-to-end validation and quality assurance for NGS.
  18. Puget et al., BRCA1 pseudogene homology and assay-design risk.
  19. MRC Holland SALSA MLPA P002 BRCA1/BRCA2 product guidance.
  20. ICR96 hereditary cancer panel CNV benchmark and MLPA truth set.
  21. European Genome-phenome Archive study EGAS00001002428 (ICR96 reads; controlled access).
  22. European Genome-phenome Archive dataset EGAD00001003335 (ICR96; DAC approval required).
  23. European Genome-phenome Archive dataset EGAD00001003400 (panelcnDataset; DAC approval required).
  24. NIST/GIAB structural-variant benchmark in dbVar.
Bottom line. Literature supports a broader normalization domain, candidate-excluded reference modeling, joint calibration, orthogonal confirmation, and physical validation. It does not support promoting the present BRCA-only count slice by choosing a more permissive threshold.