Nuclear Localization Signal Peptide: The Address Tag on a Protein
A nuclear localisation signal, usually abbreviated to NLS, is a short stretch of residues that acts as an address label on a protein, marking it for transport from the cytosol into the nucleus. It is typically fewer than about thirty residues, seldom has a stable structure of its own, and is defined by what it does rather than by a strict consensus sequence. The canonical version, taken from the SV40 large T antigen, is seven residues long: proline-lysine-lysine-lysine-arginine-lysine-valine, written PKKKRKV. Replacing a single lysine in that run with a neutral residue largely abolishes import, which is about as clear a demonstration of sequence specificity as peptide biology offers.
The pathway that reads the tag has three moving parts. Importin alpha acts as the adaptor that binds the basic motif. Importin beta carries the complex through the nuclear pore by engaging the phenylalanine-glycine repeats that line its channel. Inside the nucleus, Ran bound to GTP binds importin beta and releases the cargo, and the RanGTP gradient across the nuclear envelope therefore supplies direction rather than energy at the pore itself. This page describes the motif, compares it with other sorting signals, and covers how NLS sequences are used as laboratory tags. The chemical background sits in the peptide chemistry and structure reference, and nothing here is a protocol or any form of advice.
What Recognises the Motif
Two features define a classical NLS. It is enriched in lysine and arginine, whose side chains carry permanent positive charge, and it sits on part of the protein that stays exposed rather than buried in a fold, frequently in a flexible loop or terminal region. Two patterns are recognised. A monopartite signal is one continuous run of basic residues, as in PKKKRKV, which occupies the major binding pocket on importin alpha. A bipartite signal has two shorter basic clusters separated by a poorly conserved linker, typically around ten to twelve residues long; the classical example comes from nucleoplasmin, where the linker loops so that one cluster binds the major pocket and the second occupies a minor pocket nearby.
Importin alpha itself is built from a series of armadillo repeats forming a curved surface, and the two pockets sit within that surface. Binding is driven largely by electrostatics, since the pockets are lined with acidic residues that complement the cationic cluster, which explains why the motif tolerates substitutions so poorly at the basic positions and so well elsewhere. It also explains why phosphorylation near the signal can switch nuclear import off, since adding a negative charge next to the basic cluster disrupts binding. Some signals are not classical at all. The PY-NLS recognised by transportin contains a proline-tyrosine element within a dispersed basic and glycine-rich context, and several other import receptors exist for distinct cargo classes.
| Feature | Nuclear localisation signal | Secretory signal peptide | Mitochondrial targeting peptide |
|---|---|---|---|
| Position in the chain | Anywhere, often internal or near a terminus | N-terminus | N-terminus |
| Typical length | About 4 to 30 residues | About 15 to 30 residues | About 15 to 50 residues |
| Composition | Rich in lysine and arginine | Continuous nonpolar core preceded by positive charges | Enriched in serine, threonine and arginine, few acidic residues |
| Structure required | Usually none; often disordered | Helix-forming hydrophobic stretch | Amphipathic helix |
| Cleaved after import | No, generally retained | Yes, by signal peptidase | Yes, by matrix processing peptidase |
| Recognition machinery | Importin alpha adaptor plus importin beta | Signal recognition particle then the Sec61 channel | TOM and TIM translocase complexes |
NLS Sequences as Laboratory Tags
Because the machinery is conserved across eukaryotes, a well-characterised NLS can be appended to almost anything and expect to work. The most common application is tagging a fluorescent protein, so that a construct carrying two or three copies of PKKKRKV fused to green fluorescent protein accumulates in the nucleus and provides a clean visual marker. Similar logic applies to genome-editing nucleases and to recombinases, where nuclear entry is the first requirement for activity and adding a strong NLS improves the fraction of expressed protein that reaches the substrate. These are research techniques described in the literature, not instructions, and their efficiency depends on the host system.
Synthetic NLS peptides are also used directly. A short cationic peptide carrying the motif can be conjugated to cargo or mixed with a recombinant protein to study import kinetics, and it can serve as a competitor that occupies importin alpha and slows the uptake of natural substrates, which is a standard way to demonstrate that uptake is receptor-mediated rather than diffusive. Because these peptides are short usually between seven and twenty residues and highly charged, they dissolve readily in water and are straightforward to synthesise, though their high positive charge makes them prone to nonspecific association with surfaces and nucleic acids.
Reading and Checking an NLS Sequence
Where the material itself is concerned, verification looks much the same as for any other synthetic peptide. Because NLS peptides are short, their identity can be confirmed by measuring the intact mass and comparing it with the value computed from the sequence, while purity is checked by reversed-phase chromatography, and these two results appear together on a certificate of analysis naming the batch and method. Handling questions come down to the usual problems with cationic peptides: they can adsorb to container walls, so carrier proteins or low-binding plastics are commonly discussed in protocols, and repeated freeze-thaw of dilute solution is the usual source of loss. General conventions are collected in storage and freeze-thaw handling and identity checking in how synthetic peptides are checked.
One last point connects this topic back to the rest of peptide science. Nothing about an NLS makes it a special kind of molecule. It is residues joined by the same amide bonds running N-terminus to C-terminus, and its activity comes entirely from which residues sit where in the sequence and whether they are exposed. That is true of every short motif discussed on this site, from sorting signals to cleavage sites to adhesive recognition sequences. A handful of positions decides the behaviour of a chain that may be hundreds of residues long, which is the recurring argument for studying peptides rather than only studying whole proteins. Residue numbering and terminal naming conventions are covered in how chain ends are named and numbered.
Frequently asked questions
Is a nuclear localisation signal itself a peptide?
It is a short peptide motif within a larger protein, though synthetic peptides carrying it are sold as research reagents. The classical examples run from seven residues up to about thirty, rich in lysine and arginine, with no fixed structure required.
What is the difference between monopartite and bipartite signals?
A monopartite signal is one continuous cluster of basic residues, such as PKKKRKV, occupying the major pocket on importin alpha. A bipartite signal has two clusters separated by a linker of roughly ten to twelve residues, with each cluster engaging a separate pocket.
Does adding an NLS tag change the peptide or protein it is attached to?
It usually changes behaviour rather than chemistry. Multiple basic tags add a large cationic patch that may bind nucleic acids nonspecifically and can alter localisation or aggregation. Positioning matters too, since a buried or sterically blocked signal will not be recognised.
Related reading
Amphiphilic Peptide Meaning: One Face Oily, One Face Wet
An amphiphilic peptide keeps nonpolar residues on one face of its helix and charged or polar residues on the other. Heli
Peptide C Terminus: Structure, Charge and Modifications
The C-terminus is the end of a peptide carrying the free carboxyl group: its charge, pKa, why sequences run N to C, and
Peptide Storage Best Practices: Powder, Solution, Light and Cold Chain
Storage conventions for lyophilised powder and solution, moisture and oxygen risks, aliquoting, labelling and cold-chain
Sources & further reading
- UniProt entry for SV40 large T antigen (LT_SV40, P03070) — https://www.uniprot.org/uniprotkb/P03070/entry
- NHGRI genetics glossary: peptide — https://www.genome.gov/genetics-glossary/Peptide
- PDB-101 educational resources, RCSB — https://pdb101.rcsb.org/
This page is part of the What Peptides Are: Structure, Bonds and How Chains Are Built guide.
Questions about method, arithmetic or sourcing on this page? Message the editorial desk.
Message us