"Base" models were actually trained with some GPT instruct datasets

Wonderful_Ad_5134 · 1 year ago

Wonderful_Ad_5134 · 1 year ago

Llama2 has been pre-trained on old data (before the chatGPT AI poisoning was significant)

“Data Freshness The pretraining data has a cutoff of September 2022, but some tuning data is more recent, up to July 2023.”

“Model Dates Llama 2 was trained between January 2023 and July 2023.”

StableLM3b has been trained on more recent datasets (cutoff of march 2023) yet it doesn’t have this amount of chatgpt poisoning in it