Ensembl mobile site help

Things to know when navigating the Ensembl mobile site

Search box

Use the search box at the top right of all Ensembl views to search for a gene, phenotype, sequence variant, and more.

Top navigation

Touch MENU button to open the main menu and touch again to close.

Touch MENU

Left hand side menu

Touch the left menu icon () or swipe right to open the side menu and touch anywhere outside the menu or touch the cross icon or swipe left to close.

The ? icon

Touch the icon to get help

And don't forget to send us your comments using the feedback link inside the main menu.

EnsemblEnsembl Home

Human assembly and gene annotation

Assembly

This site provides a data set based on the December 2013 Homo sapiens high coverage assembly GRCh38 from the Genome Reference Consortium. This assembly is used by UCSC to create their hg38 database. The data set consists of gene models built from the genewise alignments of the human proteome as well as from alignments of human cDNAs using the cDNA2genome model of exonerate.

This release of the assembly has the following properties:

  • contig length total 3.4 Gb.
  • chromosome length total 3.1 Gb (excluding haplotypes).

It also includes 261 alt loci scaffolds, mainly in the LRC/KIR complex on chromosome 19 (35 alternate sequence representations) and the MHC region on chromosome 6 (7 alternate sequence representations).

Watch a video on YouTube about patches and haplotypes in the Human genome.

Patches

As the GRC maintains and improves the assembly, patches are being introduced. Currently, assembly patches are of two types:

  • Novel patch: new sequences that add alternative sequence at a loci and will remain as haplotypes in the next major assembly release by GRC
  • Fix patch: sequences that correct the reference sequence and will replace the given region of the reference assembly at the next major assembly release by GRC.

The genome assembly represented here corresponds to GenBank Assembly ID GCA_000001405.27

Gene annotation

The Ensembl human gene annotations have been updated using Ensembl's automatic annotation pipeline. The updated annotation incorporates new protein and cDNA sequences which have become publicly available since the last GRCh38 genebuild (December 2013).

In the current release, we continue to display a joint gene set based on the merge between the automatic annotation from Ensembl and the manually curated annotation from Havana. See the statistics table, right, for the corresponding GENCODE version number. The Consensus Coding Sequence (CCDS) identifiers have also been mapped to the annotations. More information about the CCDS project.

Updated manual annotation from Havana is merged into the Ensembl annotation every release. Transcripts from the two annotation sources are merged if they share the same internal exon-intron boundaries (i.e. have identical splicing pattern) with slight differences in the terminal exons allowed. Importantly, all Havana transcripts are included in the final Ensembl/Havana merged (GENCODE) gene set.

Neanderthal genome

A preliminary assembly of the Neanderthal (Homo sapiens neanderthalensis) genome is available via the Neanderthal Genome Browser, an Ensembl-powered project based at the Max Planck Institute.

More information

General information about this species can be found in Wikipedia.

Statistics

Summary

AssemblyGRCh38.p12 (Genome Reference Consortium Human Build 38), INSDC Assembly GCA_000001405.27, Dec 2013
Base Pairs3,609,003,417
Golden Path Length3,096,649,726
Annotation providerEnsembl
Annotation methodFull genebuild
Genebuild startedJan 2014
Genebuild releasedJul 2014
Genebuild last updated/patchedJul 2018
Database version94.38
Gencode versionGENCODE 29

Gene counts (Primary assembly)

Coding genes20,418 (incl 650 readthrough)
Non coding genes22,107
Small non coding genes4,871
Long non coding genes15,014 (incl 284 readthrough)
Misc non coding genes2,222
Pseudogenes15,195 (incl 8 readthrough)
Gene transcripts206,762

Gene counts (Alternative sequence)

Coding genes2,958 (incl 46 readthrough)
Non coding genes1,429
Small non coding genes278
Long non coding genes974 (incl 39 readthrough)
Misc non coding genes177
Pseudogenes1,754
Gene transcripts20,652

Other

Genscan gene predictions51,153
Short Variants665,437,818
Structural variants6,013,031

About this species