Neoantigens

Identifying novel immunological targets created by tumor-specific splicing events.

Alternative splicing can produce protein sequences the immune system has never seen. Where a novel junction falls inside a short peptide that MHC-I can display, that peptide becomes a candidate neoantigen — a tumour-specific target for T-cell recognition. SPLISOFORMS extracts these junction-spanning peptides genome-wide, scores each for MHC-I presentation and immunogenicity, and ranks them into P1–P4 priority tiers. A distinctive layer additionally flags epitopes that exist only because splicing removed a canonical regulatory (PTM) site — the resource's unique splice-enabled angle.

The Discovery Engine

From every novel splice junction to a ranked list of candidate epitopes, the pipeline runs in four stages:

  1. Junction extraction. Overlapping 8–11-mer peptides are tiled across each novel junction — novel exon, alternative splice site, intron retention, or frameshift — covering every reading frame that spans the junction.
  2. Flanking context. Each peptide is extracted with 10 residues of upstream and downstream sequence, giving the antigen-processing model the context it needs to score proteasomal cleavage right at the junction.
  3. MHC-I prediction. Every peptide × allele pair is scored by a three-tool presentation ensemble — MHCflurry 2.2, NetMHCpan-4.2, and BigMHC — plus BigMHC-IM for immunogenicity. MHCflurry additionally returns a 0–1 presentation score (proteasomal processing × MHC binding). How these combine into the P1–P4 ranking is detailed below.
  4. Candidate set. A peptide is kept if at least one tool presents it. MHCflurry binding affinity (IC50) is recorded alongside, categorised as strong (< 50 nM), weak (50–500 nM), or non-binding (dropped).

HLA Allele Panel

Predictions are run across a curated panel of 35 HLA class-I alleles selected to maximise global population coverage across HLA-A, HLA-B, and HLA-C loci. The panel follows the IEDB reference-set recommendations and adds alleles with high frequency in African and Asian populations to reduce Eurocentric bias.

HLA-A13 alleles
AlleleApprox. population frequency
HLA-A*01:01~16% European
HLA-A*02:01~28% European, ~10% Asian
HLA-A*03:01~14% European
HLA-A*11:01~22% Asian, ~5% European
HLA-A*23:01~9% African
HLA-A*24:02~18% Asian, ~5% European
HLA-A*26:01~5% European
HLA-A*29:02~6% African, ~4% European
HLA-A*30:01~8% African
HLA-A*31:01~7% Asian, ~5% European
HLA-A*32:01~6% European
HLA-A*33:01~8% Asian
HLA-A*68:01~7% African, ~3% European
HLA-B16 alleles
AlleleApprox. population frequency
HLA-B*07:02~13% European
HLA-B*08:01~10% European
HLA-B*13:02~7% Asian
HLA-B*15:01~9% European
HLA-B*15:02~8% Asian (SJS-associated)
HLA-B*15:03~10% sub-Saharan African
HLA-B*18:01~8% European
HLA-B*27:05~8% European (AS-associated)
HLA-B*35:01~8% European / Latino
HLA-B*38:01~4% European
HLA-B*40:01~9% Asian
HLA-B*44:02~10% European
HLA-B*44:03~7% African
HLA-B*51:01~7% Asian / Mediterranean
HLA-B*53:01~10% sub-Saharan African
HLA-B*58:01~8% Asian (allopurinol-HSR)
HLA-C6 alleles
AlleleApprox. population frequency
HLA-C*03:04~10% European / Asian
HLA-C*04:01~12% African / European
HLA-C*06:02~9% European
HLA-C*07:01~18% European
HLA-C*07:02~14% European
HLA-C*12:03~5% European

MHCflurry HLA-C coverage is sparser than HLA-A/B. These alleles fall back to affinity-only mode when the presentation model lacks training data.

HLA-A — 13 allelesHLA-B — 16 allelesHLA-C — 6 allelesTotal — 35 alleles

Priority Ranking (P1–P4)

Candidates are ranked on immunological evidence, not on splice/PTM origin. The headline priority tier is a transparent conjunction of two orthogonal axes: multi-tool MHC presentation consensus and BigMHC immunogenicity. Splice origin, PTM disruption, expression and fold quality are reported as decorating flags — so we can state, for example, "of P1 epitopes, X % are splice-enabled" rather than gating the tier on it.

P1 · High-confidence

Strong multi-tool presentation (≥ 2 of 3 tools call the peptide a strong binder) AND predicted immunogenic (BigMHC-IM ≥ 0.5). The candidates most likely to be presented and elicit a T-cell response.

P2 · Strong presentation

Strong multi-tool presentation, but not predicted immunogenic. Robustly presented; immunogenicity support is absent.

P3 · Supported

Moderate presentation (a single strong call, or ≥ 2 weak calls) AND predicted immunogenic.

P4 · Candidate

Presentation evidence from a single tool or without immunogenicity support. Retained as a candidate but the weakest immunological evidence.

Presentation consensus (the backbone)

Three predictors are run per peptide × allele pair — MHCflurry 2.x, NetMHCpan-4.2 (EL) and BigMHC EL — and vote at two stringencies (strong / weak) using field-standard eluted-ligand %rank cutoffs. The count of votes sets the presentation tier.

Strong

≥ 2 of 3 tools call the peptide a strong binder: MHCflurry %rank ≤ 0.5, NetMHCpan-4.2 EL %rank ≤ 0.5, or BigMHC EL ≥ 0.5.

Moderate

Exactly one strong call, or ≥ 2 tools presenting at the weak stringency (%rank ≤ 2.0 / BigMHC EL ≥ 0.25).

Weak

A single tool presents the peptide at the weak stringency.

None

No tool presents the peptide — excluded from the candidate set.

Decorating flags

Orthogonal annotations layered on top of any tier. They describe why a candidate is interesting (splice origin, PTM disruption) or add supporting evidence (expression, fold quality) — but they never change the priority tier.

Immunogenic

BigMHC-IM immunogenicity score ≥ 0.5 — predicted T-cell recognition. This is the immunogenicity axis of the P-tier and is also surfaced as a standalone flag.

Splice-enabled

The epitope exists only because of the splice event (a canonical inhibitory PTM site is lost at/near the junction). The resource's unique angle — reported, not used to rank.

Expressed

The isoform is detected in the underlying expression data (≥ 1 sample with full-length support). Currently derived from the ccRCC long-read set.

PTM-disrupting

The novel sequence ablates a canonical PTM site inside the epitope (ptm_canonical_residue_lost = TRUE).

High-fidelity fold

The epitope window is confidently modelled (avg pLDDT > 70, min pLDDT > 50) — the structural context is trustworthy.

Interpretation. All tiers are predicted, not experimentally validated. Presentation (the P-tier backbone) is well benchmarked; immunogenicity relies on a single model (BigMHC-IM) and should be read as supporting evidence. Because neoepitopes are junction-derived they are novel relative to the canonical protein, but a peptide may still occur elsewhere in the normal proteome — a proteome-wide uniqueness filter is a planned addition.

Epitope Atlas

The Epitope Atlas (accessible via the Results page toggle) provides a proteome-wide view of all predicted neoantigens in a single searchable, filterable table — complementing the per-isoform neoantigen panel.

Deduplication

Each peptide × isoform pair is reduced to a single representative row — its best-scoring allele hit — so the same junction peptide can't inflate counts across every allele in the panel.

Sortable Columns

Sort by priority tier, presentation (multi-tool consensus), immunogenicity, IC50, pLDDT, junction type, gene, or isoform. Default order is priority ascending (P1 first). Column headers show sort direction and carry inline tooltips explaining each metric.

Filters

Filter simultaneously by gene symbol, priority tier (P1–P4), decorating flag (immunogenic, splice-enabled, expressed, PTM-disrupting, high-fidelity fold), binding category, HLA allele (dynamically loaded), junction type, and minimum pLDDT.

Summary Statistics

Header chips show real-time counts per priority tier (P1–P3), immunogenic and expressed peptides, plus unique peptides and unique genes matching the active filter set.

Tools & sources