The similarity may be partly the image model's fault – especially if it's been post-trained/distilled towards performance, correctness and "quality" (for some value of "quality" anyway) which inevitably occurs at the expense of creativity and variation.
But I think it might be more about the LLM's lack of creativity in coming up with the prompt for the image (or "embellishing" a user's prompt), or directly the embedding vector if the LLM is itself the image model's text encoder. Then whatever the image gen outputs is simply an instance of the GIGO principle. It would be interesting to test whether similar cliches and motifs also occur if you ask the model to create an SVG rather than a raster image.
Its really interesting to me to read books from the 80s set in the 2000s.
The Soviet Union was there chugging along, a monolithic and unstoppable foil to Murica! Everyone things nothing ever changes until it does. I expect that to be largely the same here if it's going to happen?
That said, I doubt we get a civil war, a more effective coup maybe? Or Collapse? But a civil war seems unlikely to me. We kind of already have lived through a revolution in the last couple years but people don't really see it.
reply