cram2frags¶
At a glance
Repository: atlasxomics/cram2frag_wf · Display name: cram2frags · Modality: Epigenomics · Stage: Preprocessing
flowchart LR
C2F["cram2frags_task<br/>CRAM → fragments"]:::process
METRICS["metrics_task<br/>fragment metrics + report"]:::process
C2F --> METRICS
classDef process stroke:#818cf8,fill:#eef2ff
Workflow task DAG — an indexed CRAM is converted to a fragments file, then summarized by the metrics task.
Overview¶
cram2frags converts an indexed CRAM alignment into a tabix-indexed fragments file, producing the same analysis-ready input as ATX epigenomic preprocessing — but starting from already-aligned reads rather than raw FASTQ.
Use it when sequencing is delivered as aligned CRAM (e.g. Ultima runs) instead of FASTQ. The resulting fragments file feeds the same downstream optimization and secondary analysis Workflows.
Single-end, spatially barcoded CRAM
The input is an indexed CRAM from single-end sequencing whose reads carry
spatial barcode tags. The CRAM's .crai index is required alongside it;
a .md5 checksum and a run .json are picked up automatically when present.
Steps¶
cram2frags_task— Converts the CRAM to fragments. Splits the alignment per chromosome and converts to BAM (batchCRAM2ChrBAM.sh, parallelized acrossmax_jobs×cores_per_job), optionally dropping reads without ana3tag (filter_a3), derives each read's spatial position from its two barcodes according tobc_order, then sorts, merges, andbgzip-compresses the fragments (batchSortFrags_v4.sh) and builds a tabix index. Also generates fragment statistics and a chromosome fragment map (fragStats_sum.py,chromFragMap.py).metrics_task— Computes fragment-level QC from the fragments file (atac_analysis_summary.py) — per-tixel metrics, peak list, and QC plots — and assembles the standalone HTML report (fragReport.py). When the run's Ultima.jsonis available, it additionally renders an Ultima sequencing report.
Inputs¶
| Parameter | Type | Default | Description |
|---|---|---|---|
cram_file |
LatchFile | — | Indexed CRAM alignment from single-end sequencing. Its .crai must sit alongside it. |
project_name |
str | — | Output folder name. |
genome |
enum | hg38 |
Reference genome — hg38, mm10, or rnor6. |
bc_order |
enum | bb,ba |
Spatial barcode order — which barcode is used for rows vs. columns. |
filter_a3 |
bool | True |
Drop reads without an a3 tag. |
Compute parameters
| Parameter | Default | Description |
|---|---|---|
cores_per_job |
4 |
Cores allocated per parallel job. |
max_jobs |
10 |
Maximum jobs run in parallel. |
Outputs¶
Written to the project's cram2frags/ directory.
cram2frags/<project_name>/
├── fragments.sort.bed.gz # the fragments file
├── fragments.sort.bed.gz.tbi # tabix index
└── frag_metrics/ # QC metrics & reports
├── fragment_analysis_report.html
├── fragments_adata_obs.csv
├── fragments_peakList.bed.gz
└── ultima_report.html # when a run .json is present
| File | Description |
|---|---|
fragments.sort.bed.gz |
The fragments file — BED-like, tab-delimited, bgzip-compressed; each row is a fragment. The input to downstream Workflows. |
fragments.sort.bed.gz.tbi |
Tabix index for fast coordinate lookup. |
frag_metrics/fragment_analysis_report.html |
Self-contained HTML QC report. |
frag_metrics/fragments_adata_obs.csv |
Per-tixel (per-barcode) metrics table. |
frag_metrics/fragments_peakList.bed.gz |
Called peaks for the run. |
frag_metrics/ultima_report.html |
Ultima sequencing-run report, generated when the run's .json is available. |
Example run¶
(Representative LaunchPlan / batch-table example to be added.)