Neoantigens
Identifying novel immunological targets created by tumor-specific splicing events.
Alternative splicing can produce protein sequences the immune system has never seen. Where a novel junction falls inside a short peptide that MHC-I can display, that peptide becomes a candidate neoantigen — a tumour-specific target for T-cell recognition. SPLISOFORMS extracts these junction-spanning peptides genome-wide, scores each for MHC-I presentation and immunogenicity, and ranks them into P1–P4 priority tiers. A distinctive layer additionally flags epitopes that exist only because splicing removed a canonical regulatory (PTM) site — the resource's unique splice-enabled angle.
The Discovery Engine
From every novel splice junction to a ranked list of candidate epitopes, the pipeline runs in four stages:
- Junction extraction. Overlapping 8–11-mer peptides are tiled across each novel junction — novel exon, alternative splice site, intron retention, or frameshift — covering every reading frame that spans the junction.
- Flanking context. Each peptide is extracted with 10 residues of upstream and downstream sequence, giving the antigen-processing model the context it needs to score proteasomal cleavage right at the junction.
- MHC-I prediction. Every peptide × allele pair is scored by a three-tool presentation ensemble — MHCflurry 2.2, NetMHCpan-4.2, and BigMHC — plus BigMHC-IM for immunogenicity. MHCflurry additionally returns a 0–1 presentation score (proteasomal processing × MHC binding). How these combine into the P1–P4 ranking is detailed below.
- Candidate set. A peptide is kept if at least one tool presents it. MHCflurry binding affinity (IC50) is recorded alongside, categorised as strong (< 50 nM), weak (50–500 nM), or non-binding (dropped).
HLA Allele Panel
Predictions are run across a curated panel of 35 HLA class-I alleles selected to maximise global population coverage across HLA-A, HLA-B, and HLA-C loci. The panel follows the IEDB reference-set recommendations and adds alleles with high frequency in African and Asian populations to reduce Eurocentric bias.
| Allele | Approx. population frequency |
|---|---|
| HLA-A*01:01 | ~16% European |
| HLA-A*02:01 | ~28% European, ~10% Asian |
| HLA-A*03:01 | ~14% European |
| HLA-A*11:01 | ~22% Asian, ~5% European |
| HLA-A*23:01 | ~9% African |
| HLA-A*24:02 | ~18% Asian, ~5% European |
| HLA-A*26:01 | ~5% European |
| HLA-A*29:02 | ~6% African, ~4% European |
| HLA-A*30:01 | ~8% African |
| HLA-A*31:01 | ~7% Asian, ~5% European |
| HLA-A*32:01 | ~6% European |
| HLA-A*33:01 | ~8% Asian |
| HLA-A*68:01 | ~7% African, ~3% European |
| Allele | Approx. population frequency |
|---|---|
| HLA-B*07:02 | ~13% European |
| HLA-B*08:01 | ~10% European |
| HLA-B*13:02 | ~7% Asian |
| HLA-B*15:01 | ~9% European |
| HLA-B*15:02 | ~8% Asian (SJS-associated) |
| HLA-B*15:03 | ~10% sub-Saharan African |
| HLA-B*18:01 | ~8% European |
| HLA-B*27:05 | ~8% European (AS-associated) |
| HLA-B*35:01 | ~8% European / Latino |
| HLA-B*38:01 | ~4% European |
| HLA-B*40:01 | ~9% Asian |
| HLA-B*44:02 | ~10% European |
| HLA-B*44:03 | ~7% African |
| HLA-B*51:01 | ~7% Asian / Mediterranean |
| HLA-B*53:01 | ~10% sub-Saharan African |
| HLA-B*58:01 | ~8% Asian (allopurinol-HSR) |
| Allele | Approx. population frequency |
|---|---|
| HLA-C*03:04 | ~10% European / Asian |
| HLA-C*04:01 | ~12% African / European |
| HLA-C*06:02 | ~9% European |
| HLA-C*07:01 | ~18% European |
| HLA-C*07:02 | ~14% European |
| HLA-C*12:03 | ~5% European |
MHCflurry HLA-C coverage is sparser than HLA-A/B. These alleles fall back to affinity-only mode when the presentation model lacks training data.
Priority Ranking (P1–P4)
Candidates are ranked on immunological evidence, not on splice/PTM origin. The headline priority tier is a transparent conjunction of two orthogonal axes: multi-tool MHC presentation consensus and BigMHC immunogenicity. Splice origin, PTM disruption, expression and fold quality are reported as decorating flags — so we can state, for example, "of P1 epitopes, X % are splice-enabled" rather than gating the tier on it.
Strong multi-tool presentation (≥ 2 of 3 tools call the peptide a strong binder) AND predicted immunogenic (BigMHC-IM ≥ 0.5). The candidates most likely to be presented and elicit a T-cell response.
Strong multi-tool presentation, but not predicted immunogenic. Robustly presented; immunogenicity support is absent.
Moderate presentation (a single strong call, or ≥ 2 weak calls) AND predicted immunogenic.
Presentation evidence from a single tool or without immunogenicity support. Retained as a candidate but the weakest immunological evidence.
Presentation consensus (the backbone)
Three predictors are run per peptide × allele pair — MHCflurry 2.x, NetMHCpan-4.2 (EL) and BigMHC EL — and vote at two stringencies (strong / weak) using field-standard eluted-ligand %rank cutoffs. The count of votes sets the presentation tier.
≥ 2 of 3 tools call the peptide a strong binder: MHCflurry %rank ≤ 0.5, NetMHCpan-4.2 EL %rank ≤ 0.5, or BigMHC EL ≥ 0.5.
Exactly one strong call, or ≥ 2 tools presenting at the weak stringency (%rank ≤ 2.0 / BigMHC EL ≥ 0.25).
A single tool presents the peptide at the weak stringency.
No tool presents the peptide — excluded from the candidate set.
Decorating flags
Orthogonal annotations layered on top of any tier. They describe why a candidate is interesting (splice origin, PTM disruption) or add supporting evidence (expression, fold quality) — but they never change the priority tier.
BigMHC-IM immunogenicity score ≥ 0.5 — predicted T-cell recognition. This is the immunogenicity axis of the P-tier and is also surfaced as a standalone flag.
The epitope exists only because of the splice event (a canonical inhibitory PTM site is lost at/near the junction). The resource's unique angle — reported, not used to rank.
The isoform is detected in the underlying expression data (≥ 1 sample with full-length support). Currently derived from the ccRCC long-read set.
The novel sequence ablates a canonical PTM site inside the epitope (ptm_canonical_residue_lost = TRUE).
The epitope window is confidently modelled (avg pLDDT > 70, min pLDDT > 50) — the structural context is trustworthy.
Interpretation. All tiers are predicted, not experimentally validated. Presentation (the P-tier backbone) is well benchmarked; immunogenicity relies on a single model (BigMHC-IM) and should be read as supporting evidence. Because neoepitopes are junction-derived they are novel relative to the canonical protein, but a peptide may still occur elsewhere in the normal proteome — a proteome-wide uniqueness filter is a planned addition.
Epitope Atlas
The Epitope Atlas (accessible via the Results page toggle) provides a proteome-wide view of all predicted neoantigens in a single searchable, filterable table — complementing the per-isoform neoantigen panel.
Deduplication
Each peptide × isoform pair is reduced to a single representative row — its best-scoring allele hit — so the same junction peptide can't inflate counts across every allele in the panel.
Sortable Columns
Sort by priority tier, presentation (multi-tool consensus), immunogenicity, IC50, pLDDT, junction type, gene, or isoform. Default order is priority ascending (P1 first). Column headers show sort direction and carry inline tooltips explaining each metric.
Filters
Filter simultaneously by gene symbol, priority tier (P1–P4), decorating flag (immunogenic, splice-enabled, expressed, PTM-disrupting, high-fidelity fold), binding category, HLA allele (dynamically loaded), junction type, and minimum pLDDT.
Summary Statistics
Header chips show real-time counts per priority tier (P1–P3), immunogenic and expressed peptides, plus unique peptides and unique genes matching the active filter set.