2024-2026: AI agents, video generation (Sora, Veo), real-time generation, open-source models (Llama 3, Mistral)
Key Breakthroughs
GANs (2014) — First framework for generating realistic synthetic data through adversarial training
Transformer (2017) — Replaced RNNs with attention; enabled parallel training and much larger models
Diffusion Models (2020) — Surpassed GANs for image generation with better diversity and fidelity
RLHF (2022) — Aligned language models with human preferences through reinforcement learning
Multimodal (2023-2024) — Models that understand and generate across text, image, audio, and video
Scaling Laws
Research shows that model performance improves predictably with increases in model size, dataset size, and compute. This empirical finding drove the "bigger is better" race, though recent focus has shifted to efficiency and alignment.
Practice Task: Create a visual timeline of generative AI milestones with at least 8 key events. For each event, note the model, parameters (if applicable), and what made it significant. Research one model from the timeline in depth and write a one-page summary of its architecture and impact.