Opening LLMs for Users
codex.town
Playing and extending popular LLMs
Thank you for your donations!
codex.town
Codex Town – сообщество исследователей, практиков, экспертов, билдеров и интересующихся новыми технологиями.
codex.town
codex.town
codex.town
Codex.plan
Today:
Next time:
Introduction
codex.town
codex.town
Large Language Model
LLM is a Transformer Model trained on huge dataset
A transformer model is a neural network that learns context and thus meaning by tracking relationships in sequential data like the words in this sentence.
codex.town
Attention is all you need
Transformer models apply an evolving set of mathematical techniques, called attention or self-attention, to detect subtle ways even distant data elements in a series influence and depend on each other.
codex.town
Transformer architecture
codex.town
Transfomer based LLMs types
Review of open LLMs
codex.town
codex.town
LLMs market progress
codex.town
Llama
LLaMA (Large Language Model Meta AI) is a large language model (LLM) released by Meta AI in February 2023. Four model sizes were trained: 7, 13, 33 and 65 billion parameters. LLaMA's developers reported that the 13B parameter model's performance on most NLP benchmarks exceeded that of the much larger GPT-3 (with 175B parameters) and that the largest model was competitive with state of the art models such as PaLM
codex.town
T5
The T5 Transformer Model was introduced in 2020 by the Google AI team and stands for Text-To-Text Transfer Transformer (5 Ts, or, in our case, T5). The main problem T5 addresses is the lack of systematic studies comparing best practices in the field of NLP.
T5 is an encoder-decoder model pre-trained on a multi-task mixture of unsupervised and supervised tasks and for which each task is converted into a text-to-text format.
codex.town
Vicuna
An open-source chatbot trained by fine-tuning LLaMA on user-shared conversations collected from ShareGPT. Preliminary evaluation using GPT-4 as a judge shows Vicuna-13B achieves more than 90%* quality of OpenAI ChatGPT and Google Bard while outperforming other models like LLaMA and Stanford Alpaca in more than 90%* of cases. The cost of training Vicuna-13B is around $300. The code and weights, along with an online demo, are publicly available for non-commercial use.
codex.town
Falcon
Falcon is a new family of state-of-the-art language models created by the Technology Innovation Institute in Abu Dhabi, and released under the Apache 2.0 license. Notably, Falcon-40B is the first “truly open” model with capabilities rivaling many current closed-source models. This is fantastic news for practitioners, enthusiasts, and industry, as it opens the door for many exciting use cases.
The Falcon family is composed of two base models: Falcon-40B and its little brother Falcon-7B. The 40B parameter model currently tops the charts of the Open LLM Leaderboard, while the 7B model is the best in its weight class.
codex.town
Open-Assistant
OpenAssistant is an artificial intelligence (AI) open source chat-based assistant that understands tasks, can interact with third-party systems and retrieve information dynamically to do so. The project is developed by a group of volunteers in collaboration with LAION. One of the goals for development includes free access to large language models that can be run locally on consumer hardware. The project is backed by a worldwide crowdsourcing effort involving over 13,500 volunteers who have created 600k human-generated data points
codex.town
BLOOM
BigScience Large Open-science Open-access Multilingual Language Model (BLOOM) is a transformer-based large language model. It was created by over 1000 AI researchers to provide a free large language model for everyone who wants to try. Trained on around 366 billion tokens over March through July 2022, it is considered an alternative to OpenAI's GPT-3 with its 176 billion parameters.
Review of UI Tools for LLM
codex.town
codex.town
UI Tools for LLMs
Dive into LLM Fine-Tuning Techniques
codex.town
codex.town
Fine-Tuning Techniques
codex.town
Full fine-tuning
Instruction fine-tuning, where all of the model's weights are updated is known as full fine-tuning. The process results in a new version of the model with updated weights. It is important to note that just like pre-training, full fine tuning requires enough memory and compute budget to store and process all the gradients, optimizers and other components that are being updated during training.
codex.town
Parameter efficient fine-tuning
In contrast to full fine-tuning where every model weight is updated during supervised learning, parameter efficient fine tuning methods only update a small subset of parameters
Low-rank Adaptation, or LoRA for short, is a parameter-efficient fine-tuning technique that falls into the re-parameterization category.
With prompt tuning, you add additional trainable tokens to your prompt and leave it up to the supervised learning process to determine their optimal values. The set of trainable tokens is called a soft prompt, and it gets prepended to embedding vectors that represent your input text.
codex.town
Parameter efficient fine-tuning
codex.town
Reinforcement learning by human feedback
RLHF uses reinforcement learning, or RL for short, to finetune the LLM with human feedback data, resulting in a model that is better aligned with human preferences. You can use RLHF to make sure that your model produces outputs that maximize usefulness and relevance to the input prompt. Perhaps most importantly, RLHF can help minimize the potential for harm. You can train your model to give caveats that acknowledge their limitations and to avoid toxic language and topics.
Review of LoRA and RLHF Techniques
codex.town
codex.town
Fine-tuning Llama with LoRA (try 1)
codex.town
Fine-tuning Llama with LoRA (try 2)
Conclusion and Q&A
codex.town
codex.town
Conclusion
Спасибо!
ТелеграмYoutube:
Site:
Сайт: