hdallatorre commited on
Commit
ddcbf81
1 Parent(s): e9599ba

feat: Add model card

Browse files
Files changed (1) hide show
  1. README.md +107 -0
README.md ADDED
@@ -0,0 +1,107 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: cc-by-nc-sa-4.0
3
+ widget:
4
+ - text: ACCTGA<mask>TTCTGAGTC
5
+ tags:
6
+ - DNA
7
+ - biology
8
+ - genomics
9
+ datasets:
10
+ - InstaDeepAI/multi_species_genome
11
+ - InstaDeepAI/nucleotide_transformer_downstream_tasks
12
+ ---
13
+ # nucleotide-transformer-v2-100m-multi-species
14
+
15
+ The Nucleotide Transformers are a collection of foundational language models that were pre-trained on DNA sequences from whole-genomes. Compared to other approaches, our models do not only integrate information from single reference genomes, but leverage DNA sequences from over 3,200 diverse human genomes, as well as 850 genomes from a wide range of species, including model and non-model organisms. Through robust and extensive evaluation, we show that these large models provide extremely accurate molecular phenotype prediction compared to existing methods
16
+
17
+ Part of this collection is the **nucleotide-transformer-v2-100m-multi-species**, a 100m parameters transformer pre-trained on a collection of 850 genomes from a wide range of species, including model and non-model organisms.
18
+
19
+ **Developed by:** InstaDeep, NVIDIA and TUM
20
+
21
+ ### Model Sources
22
+
23
+ <!-- Provide the basic links for the model. -->
24
+
25
+ - **Repository:** [Nucleotide Transformer](https://github.com/instadeepai/nucleotide-transformer)
26
+ - **Paper:** [The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics](https://www.biorxiv.org/content/10.1101/2023.01.11.523679v1)
27
+
28
+ ### How to use
29
+
30
+ <!-- Need to adapt this section to our model. Need to figure out how to load the models from huggingface and do inference on them -->
31
+ Until its next release, the `transformers` library needs to be installed from source with the following command in order to use the models:
32
+ ```bash
33
+ pip install --upgrade git+https://github.com/huggingface/transformers.git
34
+ ```
35
+
36
+ A small snippet of code is given here in order to retrieve both logits and embeddings from a dummy DNA sequence.
37
+ ```python
38
+ from transformers import AutoTokenizer, AutoModelForMaskedLM
39
+ import torch
40
+
41
+ # Import the tokenizer and the model
42
+ tokenizer = AutoTokenizer.from_pretrained("InstaDeepAI/nucleotide-transformer-v2-100m-multi-species")
43
+ model = AutoModelForMaskedLM.from_pretrained("InstaDeepAI/nucleotide-transformer-v2-100m-multi-species")
44
+
45
+ # Create a dummy dna sequence and tokenize it
46
+ sequences = ['ATTCTG' * 9]
47
+ tokens_ids = tokenizer.batch_encode_plus(sequences, return_tensors="pt")["input_ids"]
48
+
49
+ # Compute the embeddings
50
+ attention_mask = tokens_ids != tokenizer.pad_token_id
51
+ torch_outs = model(
52
+ tokens_ids,
53
+ attention_mask=attention_mask,
54
+ encoder_attention_mask=attention_mask,
55
+ output_hidden_states=True
56
+ )
57
+
58
+ # Compute sequences embeddings
59
+ embeddings = torch_outs['hidden_states'][-1].detach().numpy()
60
+ print(f"Embeddings shape: {embeddings.shape}")
61
+ print(f"Embeddings per token: {embeddings}")
62
+
63
+ # Compute mean embeddings per sequence
64
+ mean_sequence_embeddings = torch.sum(attention_mask.unsqueeze(-1)*embeddings, axis=-2)/torch.sum(attention_mask, axis=-1)
65
+ print(f"Mean sequence embeddings: {mean_sequence_embeddings}")
66
+ ```
67
+
68
+
69
+ ## Training data
70
+
71
+ The **nucleotide-transformer-v2-100m-multi-species** model was pretrained on a total of 850 genomes downloaded from [NCBI](https://www.ncbi.nlm.nih.gov/). Plants and viruses are not included in these genomes, as their regulatory elements differ from those of interest in the paper's tasks. Some heavily studied model organisms were picked to be included in the collection of genomes, which represents a total of 174B nucleotides, i.e roughly 29B tokens. The data has been released as a HuggingFace dataset [here](https://huggingface.co/datasets/InstaDeepAI/multi_species_genomes).
72
+
73
+ ## Training procedure
74
+
75
+ ### Preprocessing
76
+
77
+ The DNA sequences are tokenized using the Nucleotide Transformer Tokenizer, which tokenizes sequences as 6-mers tokenizer when possible, otherwise tokenizing each nucleotide separately as described in the [Tokenization](https://github.com/instadeepai/nucleotide-transformer#tokenization-abc) section of the associated repository. This tokenizer has a vocabulary size of 4105. The inputs of the model are then of the form:
78
+
79
+ ```
80
+ <CLS> <ACGTGT> <ACGTGC> <ACGGAC> <GACTAG> <TCAGCA>
81
+ ```
82
+
83
+ The tokenized sequence have a maximum length of 1,000.
84
+
85
+ The masking procedure used is the standard one for Bert-style training:
86
+ - 15% of the tokens are masked.
87
+ - In 80% of the cases, the masked tokens are replaced by `[MASK]`.
88
+ - In 10% of the cases, the masked tokens are replaced by a random token (different) from the one they replace.
89
+ - In the 10% remaining cases, the masked tokens are left as is.
90
+
91
+ ### Pretraining
92
+
93
+ The model was trained with 8 A100 80GB on 300B tokens, with an effective batch size of 1M tokens. The sequence length used was 1000 tokens. The Adam optimizer [38] was used with a learning rate schedule, and standard values for exponential decay rates and epsilon constants, β1 = 0.9, β2 = 0.999 and ε=1e-8. During a first warmup period, the learning rate was increased linearly between 5e-5 and 1e-4 over 16k steps before decreasing following a square root decay until the end of training.
94
+
95
+
96
+ ### BibTeX entry and citation info
97
+
98
+ ```bibtex
99
+ @article{dalla2023nucleotide,
100
+ title={The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics},
101
+ author={Dalla-Torre, Hugo and Gonzalez, Liam and Mendoza Revilla, Javier and Lopez Carranza, Nicolas and Henryk Grywaczewski, Adam and Oteri, Francesco and Dallago, Christian and Trop, Evan and Sirelkhatim, Hassan and Richard, Guillaume and others},
102
+ journal={bioRxiv},
103
+ pages={2023--01},
104
+ year={2023},
105
+ publisher={Cold Spring Harbor Laboratory}
106
+ }
107
+ ```