Gene3D: comprehensive structural and functional annotation of genomes.
File(s)
Author(s)
Type
Journal Article
Abstract
Gene3D provides comprehensive structural and functional annotation of most available protein sequences, including the UniProt, RefSeq and Integr8 resources. The main structural annotation is generated through scanning these sequences against the CATH structural domain database profile-HMM library. CATH is a database of manually derived PDB-based structural domains, placed within a hierarchy reflecting topology, homology and conservation and is able to infer more ancient and divergent homology relationships than sequence-based approaches. This data is supplemented with Pfam-A, other non-domain structural predictions (i.e. coiled coils) and experimental data from UniProt. In order to enhance the investigations possible with this data, we have also incorporated a variety of protein annotation resources, including protein-protein interaction data, GO functional assignments, KEGG pathways, FUNCAT functional descriptions and links to microarray expression data. All of this data can be accessed through a newly re-designed website that has a focus on flexibility and clarity, with searches that can be restricted to a single genome or across the entire sequence database. Currently Gene3D contains over 3.5 million domain assignments for nearly 5 million proteins including 527 completed genomes. This is available at: http://gene3d.biochem.ucl.ac.uk/
Date Issued
2008-01-01
Date Acceptance
2007-10-27
Citation
Nucleic Acids Research, 2008, 36 (suppl_1), pp.D414-D418
ISSN
0305-1048
Publisher
Oxford University Press
Start Page
D414
End Page
D418
Journal / Book Title
Nucleic Acids Research
Volume
36
Issue
suppl_1
Copyright Statement
© 2007 The Author(s). This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.or
g/licenses/
by-nc/2.0/uk/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.
g/licenses/
by-nc/2.0/uk/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.
Identifier
https://www.ncbi.nlm.nih.gov/pubmed/18032434
PII: gkm1019
Subjects
Databases, Protein
Evolution, Molecular
Genome, Archaeal
Genome, Bacterial
Genomics
Internet
Protein Structure, Tertiary
Proteins
Sequence Analysis, Protein
Sequence Homology, Amino Acid
User-Computer Interface
Publication Status
Published
Coverage Spatial
England
Date Publish Online
2007-11-21