AI 炒作反思:LLM 批判性阅读清单
精选资源帮助理解 AI 行业的真实现状。培养程序员的理性思维。
精选资源帮助理解 AI 行业的真实现状。培养程序员的理性思维。
添加合理的链接和对事物运作方式的良好解释。尽量避免炒作和厂商内容。热切寻求关于生产环境模型的实际一手账述。
The Illustrated Word2vec - A Gentle Intro to Word Embeddings in Machine Learning (YouTube)
Transformers as Support Vector Machines
Deep Learning Systems
Fundamental ML Reading List
Concepts from Operating Systems that Found their way into LLMS
Language Modeling is Compression
Vector Search - Long-Term Memory in AI
Eight things to know about large language models
The Scaling Hypothesis
Attention is all you Need
Scaling Laws for Neural Language Models
GPT-2: Language Models are Unsupervised Multi-Task Learners
InstructGPT: Training Language Models to Follow Instructions
GPT-3: Language Models are Few-Shot Learners
Transformers from Scratch
Five Years of GPT Progress
Lost in the Middle: How Language Models Use Long Contexts
Self-attention and transformer networks
Understanding and Coding the Attention Mechanism
Keys, Queries, and Values
What is ChatGPT doing and why does it work
My own notes from a few months back.
Karpathy's The State of GPT (YouTube)
Catching up on the weird world of LLMS
How open are open architectures?
Building an LLM from Scratch
Large Language Models in 2023 and Slides
Timeline of Transformer Models
Large Language Model Evolutionary Tree
What's in my Big Data
"The "it" in AI models is the dataset."
Extracting Training Data from ChatGPT
Why host your own LLM?
How to train your own LLMs
Hugging Face Resources on Training Your Own
Training Compute-Optimal Large Language Models
Supervised Fine-tuning
How Abilities in LLMs Are Affected by SFT
Instruction-tuning for LLMs: Survey
Direct Preference Optimization: Your Language Model is Secretly a Reward Model
RLHF and DPO Compared
The Complete Guide to LLM Fine-tuning
LLaMAntino: LLaMA 2 Models for Effective Text Generation in Italian Language - Really great overview of SOTA fine-tuning techniques
On the Structural Pruning of Large Language Models
A Gentle Introduction to 8-bit matrix multiplication
Which Quantization Method is Right for You?
Survey of Quantization for Inference
Fine-tuning with LoRA and QLoRA
Motivation for Parameter-Efficient Fine-tuning
How is LlamaCPP Possible?
How to beat GPT-4 with a 13-B Model
Efficient LLM Inference on CPUs
Tiny Language Models Come of Age
Efficiency LLM Spectrum
Building LLM Applications for Production
Challenges and Applications of Large Language Models
All the Hard Stuff Nobody talks about when building products with LLMs
Scaling Kubernetes to run ChatGPT
Numbers every LLM Developer should know
Against LLM Maximalism
A Guide to Inference and Performance
(InThe)WildChat: 570K ChatGPT Interaction Logs In The Wild
The State of Production LLMs in 2023
Machine Learning Engineering for successful training of large language models and multi-modal models.
Fine-tuning RedPajama on Slack Data
LLM Inference and K-V Cache
LLM Inference Performance Engineering: Best Practices
How to Make LLMs go Fast
Transformer Inference Arithmetic
Which serving technology to use for LLMs?
Speeding up the K-V cache
Large Transformer Model Inference Optimization
On Prompt Engineering
Prompt Engineering Versus Blind Prompting
Building RAG-Based Applications for Production
Full Fine-Tuning, PEFT, or RAG?
Prompt Engineering Guide
The Best GPUS for Deep Learning 2023
Making Deep Learning Go Brr from First Principles
Everything about Distributed Training and Efficient Finetuning
Training LLMs at Scale with AMD MI250 GPUs
ChatGPT: Jack of All Trades, Master of None
What's Going on with the Open LLM Leaderboard
Challenges in Evaluating AI Systems
LLM Evaluation Papers
Evaluating LLMs is a MineField
Generative Interfaces Beyond Chat (YouTube)
Why Chatbots are not the Future
The Future of Search is Boutique
As a Large Language Model, I
Natural Language is an Unnatural Interface
感谢所有在 Twitter、Mastodon 和 Bluesky 上提供建议的人。