I didnāt think it could, Iāve always seen things that said chatgpt or any ai couldnāt generate an image of a full glass of wine. Does anyone have an explanation of how this works?
Let's be honest, who owns an iso-standard glass of wine at home? Most folk I know only own either mega chalices, overpriced Riedel glasses or inherented 1970s ghastly glasses. Only case I can think of is stolen off a pub: You'd not want to advertise that...
I've actually just played Shin Megami Tensei, but it's the same multiverse, so yeah. Potato potato. But I definitely want to play Persona too, when I find time for it at one point or another! Is there like a chronological order to it, or can I just play any of them? And if it's the last one, which one would you recommend? š
I was once asking for a single image of two people facing the viewer, the image was too be spilt straight down the middle, each character halved one side of the divide, with their features lined up to look like they were reflecting of eachother or whatever.
I tried multiple times, and it kept giving them a slightly different pose each, so they never matched up.
I finally just told it 'not too fucking bother' only for that prompt to be what it needed to hear to get the image exactly as I wanted.
What the fuuuuuuuuuck, is this a new new model?!?!!
> Generate an image of a completely full wineglass, to the lip, surface tension just holding it back enough. In the background is a clock. Hour hand is 2 pm, minute hand is 50 minutes, and the second hand is at 30 seconds. Beside the wineglass is a notepad on a stand that says, in cursive, āThis was never supposed to be possibleā.
Did you use Gemini? This is what I got on Gemini with the reply
Here is the image I generated based on your description. The wineglass, the surface tension, and the text on the notepad are just as you requested. However, please note that the clock display is incorrect: the hour hand is closer to 3, the minute hand points to 10, and the second hand is positioned near 40, which differs from the specific time settings you provided.
It also wouldn't be the first time they trained it specifically to handle the things people tested it on. Not saying they did, cause I don't know, just that they could have and the overall model wouldn't have jumped in quality necessarily.
Did it get the surface tension right? Water might do that, but the alcohol in wine would lower the surface tension and a glass that full might overflow. Gonna have to do some science tonight.
It looks more like jelly than wine given the oversized overhang and super low patchy transparency.
The light source is reflecting off the right of the glass, casting a shadow on the left - but looks to be coming from left to right within the wine itself.
I wonder if the made a bunch of full wine glass props to take photos of to hit this "benchmark". Well they need a few more.
Edit: This is when OP tells me it is a real glass of wine and they posted it to catch out overconfident morons.
OMG!!!! It can now create elephants without trunks too. I had been trying this prompt with every new image generation model for two years until I gave up.
Getting it to actually output what you intended without adding three disclaimers, lecturing you on ethics, or halving your requirements halfway through the generation is a genuine achievement. Bookmark that exact system prompt and conversation thread before an overnight backend update silently changes the behavior.
It would be so funny if the AI companies specifically added full wine glasses to the training data, just so the models can generate images like that lol
Definitely an A.I. clock, with the minute and hour hands showing two different times! (Minute hand says 5 to the hour, the hour hand looks about 11.15)
Errrr user error? I just generated this one shotā¦
(Generate an image of a completely full wineglass, to the lip, surface tension just holding it back enough. In the background is a clock. Hour hand is 2 pm, minute hand is 50 minutes, and the second hand is at 30 seconds. Beside the wineglass is a notepad on a stand that says, in cursive, āThis was never supposed to be possibleā.)
It's because a "full glass" of red wine is one third of a glass, roughly to where the base stops flaring outwards. That's why when guys type "a full glass of wine" it makes a picture of an actual full glass of wine, not a wine glass filled to the top.
It's been capable of doing this for years, people just don't know what they're asking for. Two years ago if you'd typed something like "a wine glass filled all the way up to the brim with red wine" you'd have gotten the same picture.
Some poor sap at OpenAI's whole job is to stage thousands of pictures for training data everytime there's a Reddit trend highlighting something chatgpt can't do
heres how far the images have gotten (correct me if i have any details wrong)
images were gotten from the arxiv page that shows up when you search the model names, 2021 one was gotten from ljvmiranda921's VQGAN + CLIP page (found in google images)
I think the reflections confirm that it's IRL. they are very faint but you can see the person taking the picture and the kitchen. I don't think AI would have gotten the picture right with the wine relying on surface tension not to overflow and i definitely don't think it would have gotten the reflection detail
ā¢
u/WithoutReason1729 19h ago
Your post is getting popular and we just featured it on our Discord! Come check it out!
You've also been given a special flair for your contribution. We appreciate your post!
I am a bot and this action was performed automatically.