Function Callings - Can be connected to text external tools or API. Could be for your own database for user data. Figma plugin to create designs for you. We need to define tools and functions. Later you call the functions!
Ian Goodfellow, creator of the GANs? 2014 first time that GAN could generate a face. In 2018 was already pretty good. After that it got better and better.
AI capabilities and computational efficiency are doubling every 3 to 7 months?
AttnGAN 2017 - one of the first text to image models. Attention GAN. First time that as you type generates the 256 x 256 image.
DALL-E (2021) and DALL-E 2 (2022) - also text to image generation by open AI. They also did CLIP: connecting text and images.
Stable Diffusion 2022 - base for runways model? You generate an image by pure noise, by removing a little of the noise of the previous one on each new generation. High-Resolution Image Synthesis with Latent Diffusion Models. In Germany. They also created LAION, if breakthrough means good dataset.
Gemini Nano Banana - Combine photos, text to image, image references well taken. The 2 and the Pro are more updated versions. Can do text seamlessly.
Ian Goodfellow is the creator through his paper.
Generator - In charge to generate image. At first is so bad and discriminator would be, this is so fake. In the end it will generate better images thanks to the discriminator. Artist
Discriminator - is this image real or fake?. Artist critique
Collect lots of images. Then little bit noise and a little image and then denoise back. During training we add more noise to an image and then we ask the model to denoise to get to he clear image. It does it from steps, takes image and denoise and then again a little bit more. They’re slow because the models have to run for a variety of steps.
Autoregressive - instead of giving out 250 to denoise at once, runway only takes 4 frames and denoise and generate that part. They can achieve real time interactions thanks to this. Starting from the starting image and once generate, generates 4 frames and get the last image and do some more little by little. But because is so fast it seems like a livestream.
Gemini’s video model is really good. Kling too. Fal.ai too to use the different models. Seedance 2.0.