What happened
Clinical interpretation of genetic variation depends on the diploid genotype, including zygosity, allele dosage and whether multiple variants occur in cis on the same homologue or in trans on different homologues. Most genomic language models process haploid sequences or combine independently encoded haplotypes downstream, so they do not directly represent the paired genotype in a single sequence. We introduce a reference-aligned diploid encoding for single-nucleotide variants (SNVs) and short insertions and deletions (indels), together with unphased and phase-retaining tokenizers that accept phased genotypes and convert them to single-sequence diploid representation. Using Nucleotide Transformer v3 backbones, we continue training 8-million- and 100-million-parameter models and evaluate an auxiliary Contrastive Phase Loss (CPL) designed to retain the phasing information of the variants in contextual representations. We evaluate on a novel compound-heterozygous benchmark containing 9,460 examples. Models whose inputs did not distinguish relative phase remained near chance, whereas our diploidic models improved discrimination with AUROC 0.649, compared to 0.506 for the vocabulary-adapted control. These findings establish a method for making diploid genotype information accessible to genomic language models, rather than a universal improvement in variant prediction; validation in naturally observed, accurately phased clinical cohorts remains necessary.
