And that's it from Glasgow! It was a fun conference, with a lively community & interesting presentations. While games are not common at ACII, still, there were interesting applications in the performing arts!
Thank you to all collaborators at @InDigitalGames for their hard work!
By rewarding the generator when it matches an "ideal" arousal curve, the generator learns what to place in different parts of the level. A level that is constantly non-arousing is boring, but so is one constantly arousing. Pacing can be a reward.
Paper: matthewbarthet.com/files/PCG…
The generator puts down one track segment at a time, and repairs the track after some actions (making a loop). Then, a simple AI agent plays the track: based on the AI game-states, "arousal" is measured in every part of the track from human arousal labels in similar game-states.
Finally, Matthew Barthet closed the conference on the last day by presenting how we can generate racing tracks that would increase a player's arousal (or follow a specific arousal progression). Again, a corpus of annotated human gameplay sessions was used to "assess" arousal.
We tried to predict engagement in an unseen video of the same game as the game we trained the model on, annotated by unseen participants. Results were mixed: some games were easier to predict than others. Battle Royale games were especially hard.
Paper: antoniosliapis.com/papers/va…
We modeled engagement in an ordinal fashion, trying to predict if engagement would increase or decrease in the next time window of this video. We used pre-trained computer vision and audio models and fused them, finally training an SVM on the annotators' consensus on engagement.
At the main conference, Kosmas Pinitas presented his work on predicting viewers' engagement from gameplay videos and audio alone (without face cams or sensors). To do this, 20 participants annotated their engagement when watching 2 hours of First Person Shooter gameplay videos.
As expected, maximizing arousal (predicted or not) didn't work well. When combined with in-game rewards it 𝘴𝘰𝘮𝘦𝘵𝘪𝘮𝘦𝘴 did 𝘰𝘬. Pushing the agent to explore more states (not get stuck in one "arousing" game-state) is an important next step.
Paper: antoniosliapis.com/papers/af…
We can use this predicted emotion (based on humans' reported arousal at similar game-states) as a reward for a reinforcement learning agent. While this doesn't work alone, we can combine it with an in-game reward (such as going quickly in the right direction for a racing game).
The 3 games feature short game sessions (max 2 minutes) but have over 120 sessions from different players who annotated their arousal levels while playing (on a recording of their play). We match game-state and arousal to find how an AI agent would "feel" in a similar game-state.
In the same workshop, Matthew Barthet presented the Affectively framework, a Gym environment for making game-playing agents in 3 games. The unique selling point is that Affectively includes approximations of "agent" emotion based on a human corpus of annotated game sessions.
There's a few caveats: we don't know the "ground truth" of emotions and rely only on how they are expressed (such as loud screams). Streamers also over-emote. Finally, we used pre-trained text/audio/facecam affect models not tailored to this task.
Paper: antoniosliapis.com/papers/th…
In our experiment we looked at whether affect manifestation levels changed when entering a new room with different design features. In all our experiments, cutscenes and scripted events were by far the best predictors of increased fear, surprise and arousal. Makes sense!
We measured emotions directly from the streamers' affect manifestations (facial expressions, voice levels, utterances) and passed each through pre-trained AI models of affect to find fear, surprise or arousal levels. For this study we assume that such models are robust enough.
First off, at the Dungeons, Neurons, and Dialogues workshop, we presented "The Scream Stream: Multimodal Affect Analysis of Horror Game Spaces" where we looked at Let's Play videos of the Outlast horror game and the impact of level design choices on the streamers' emotions.
Last week I was in Glasgow for the International Conference on Affective Computing & Intelligent Interaction @acii_conf and since we had quite a few things there, let me unfurl them in a thread.
We're iterating on our CrawLLM game generator, and we'd like to know how the generated visuals and generated text is consistent with an overall theme. Fill in our 15-minute (anonymous) questionnaire (very light, with plenty of images) here!
docs.google.com/forms/d/e/1F…
Hello world, would you help out with my research? Please fill in my survey here: forms.gle/CRywoErJ9UEyWxdaA It's a 'fun' quiz about AI-generated images and text. Participation is optional and completely anonymous. The survey takes between 10 and 15 min to complete. Thank you!!
We're happy to announce a Special Issue on Large Language Models and Games at the IEEE Transactions on Games (Impact Factor: 1.7). We welcome research papers, real-world applications, opinion papers, and surveys.
🌐Details: transactions.games/special-i…
📅Deadline: 1st December 2024
If you're interested in this work, follow @MEliTAbot which is outputting the generated games (description and art) using 7 (generated) game titles as seeds.
And if you're in Aberystwyth this week for the @EvostarConf come say hi to Marvin and myself!
This research was spearheaded by Marvin Zammit, with some help from myself and @yannakakis from the @InDigitalGames
We also thank the @ai4mediaproject for supporting this and our other research on quality-diversity evolutionary search in the field of media.
Our experiments comparing MEliTA with MAP Elites (that treats image+text as a single entity) show that MEliTa can create more coherent individuals (high CLIP score) but covers less of the design space as more content is copied over (e.g. some games end up sharing the same art).
Speaking of characterization, we match cover images based on colorfulness and complexity. For game descriptions, we assign our generated description to one of 16 topics found via Latent Dirichlet allocation on the entire Steam dataset of game descriptions.
To produce new variations for the descriptions or the cover art, we implemented custom mutation operators that first destroy parts of the previous artifact and then use the trained models (GPT-2 or StableDiffusion) to repair whatever is left.
This is difficult to explain in a vacuum, so we implemented MEliTA to generate hypothetical games, specifically their title, description & cover art (using GPT-2 & StableDiffusion). These could inspire new (human) projects, but we are obviously not automating game development.
While MAP-Elites would treat the combined modalities (e.g. images and text) as a single artifact, our MEliTA algorithm produces a variation of one modality and then matches it with the most appropriate version of the other modality in the archive to create a new pairing.
This Friday at @EvostarConf, Marvin Zammit is presenting our paper on MEliTA, a variant of the MAP-Elites evolutionary algorithm that is suited for multimodal creative problems (such as simultaneous image and text generation).
Paper: antoniosliapis.com/papers/ma…
A potential solution to this caveat of finding appropriate dimensions for MAP-Elites would be to include designer models, thus learning how designers choose & adapting the features accordingly. These are ideas we can explore in future work, building on our other research.
Results are promising, especially considering the tastes of our (artificial) users with controllable selection criteria (USC). Admittedly, the system works best when the behavior characterizations/dimensions of the feature map somehow match user aesthetics & are not orthogonal.
We tested this on an architectural layout task, based on the H2020 EU project PrismArch where quickly exploring the generative space of floorplans in VR was critical. For this problem, we integrated constrained optimization using dual archives for feasible & infeasible solutions.
Evolution is also faster, contained to a much smaller selection window compared to the whole feature map. The rest of the feature map may still be populated by offspring that don't fit the window (they are still placed anywhere) but its coverage is not as broad as MAP-Elites.
We tested many ways of selecting what to show the user, to capture good parts of the space they are interested in, taking advantage of the illumination capacity of MAP-Elites. This way, the user is not overwhelmed with too many "samey" choices that cause fatigue.
Our algorithm, User Controllable MAP-Elites (UC-ME), repeats this:
1. All parents are picked from a selection window of n by n cells in the MAP-Elites feature map.
2. User chooses one favorite from 4 alternatives in this window.
3. The center of the window moves to user's choice.
Our motivation goes like this:
+ Interactive Evolution allows users to control search within the generative space, but leads to fatigue.
+ MAP-Elites & similar algorithms illuminate the whole space but show users too much data.
+ Controlling MAP-Elites search/budget solves both.
You can read the original paper and follow-up work by my lab here: antoniosliapis.com/projects/…
If you still have Java, the executable at sentientsketchbook.com/downl… should still be working!
Happy 10th anniversary Sentient Sketchbook!
The FDG 2013 paper wasn't the only one that covered Sentient Sketchbook, but it was the first one. Future papers focused more on theoretical underpinnings of such mixed-initiative co-creativity, on personalizing AI suggestions, on expanding the premise to other genres, and more.
The system "worked" because the designer hadn't invested too much time making these low-resolution "map sketches", so they were amenable to changing their design based on AI feedback. Also the small maps were much easier for the AI to evolve & optimize in (essentially) real-time.
Sentient Sketchbook was a computer-aided tool for level design of small, low-effort, level drafts. An AI provided feedback in terms of relevant playability metrics, but it also showed its own alternative level drafts at all times (not on-demand) to the designer.
User fatigue is not only an issue of Picbreeder, but rather a staple challenge within Interactive Evolution. I'd suggest the paper of Takagi on "Interactive Evolutionary Computation" from 2001, and subsequent surveys by the same author on the topic.
The last day of @FDGconf starts with an impromptu panel with @annetropy @gillianmsmith Rafa Bidarra & @TheJimWhitehead on where we are as a community, where we're going, and some life lessons for young academics. Chaired by @mtrc !
Lots of things happening at @FDGconf but I want to highlight the excellent keynote by @nielsen_holly discussing tabletop games (focusing on the UK context), their politics & more. I highlight 1 slide, emphasizing the "remember, it's just a game..." in this designer's description.
At #FDG23 the Tabletop Games Workshop is jamming! Four groups are (each) prototyping a board game in 1.5 hours, then showing it to the other groups for feedback! Micael Sousa doing a great work keeping people energized and creative!
This work originates from the working group on the same topic at the 2023 Schloss Dagstuhl Research Meeting on Game Measures & Player Experience. It was a great opportunity to give more shape to the idea, and hopefully we can work on a prototype soon as well!
In the paper, we discuss how playfulness can be introduced in the task's framing (promoting human playfulness), in the interaction (e.g. with an AI misinterpreting a human co-creator's work or intent), and in the AI itself (e.g. driven by autonomous goal-seeking or curiosity).
I'm in Lisbon this week attending the #FDG23 conference. This year sees @fdgconf getting more "physical" again, with 120+ attendees in Lisbon, 33 full & 17 short papers presented on-site. There's also 6 workshops, demos & a doctoral consortium.
More at fdg2023.org/
Accuracies are not very high, even with the best models. This is in part due to the noisy webcam data, as we wanted players to act natural. The dataset & 1st findings are promising though! In future work, we also want to look at the gameplay footage as another tension predictor.
The main takeaway of this research is that, when using facial action units (brow raise, etc.) obtained by facial detection software to predict players' tension, the data only from the opposing player can be an equally powerful predictor. Also, using both players' data works best.
We wanted users' interactions to be as natural as possible. Few constraints were placed, except to leave the room when not competing. For the same reason, players did not annotate how they were feeling; instead, this was done on the recordings by 2 experts using the PAGAN tool.
The players (17 in total) were competing for awards for the Top 3 spots. They played in a university lab, on opposing sides of 5 & 5 desktop PCs that were recording their faces via webcam & the game state via screen capture software. After cleanup, 156 videos were collected.
Next week at #FDG23 Paris Mavromoustakos-Blom is presenting our work on predicting the tension of opposing players competing in a Hearthstone tournament. In this case, the players were literally on opposite sides of a desk & could see each other.
🗒️ antoniosliapis.com/papers/mu…
But for now, I'm looking forward to hearing the opinions of other #FDG23 attendees next week. This paper was a lot of fun to write, as was this thread. If you're interested to join this PX evaluation quest, do reach out on DMs or e-mail!
📖 also in HTML: antoniosliapis.com/articles/…
This first step, based on a literature review and our own intuitions as TTRPG enthusiasts, will need to be substantiated via interviews with GMs & game scholars, and follow through a many-step questionnaire development process. This is only the start of an exciting adventure!
We're not the 1st to suggest this (although the only TTRPG PX questionnaire we found is from 2008, for one study). Our paper has 117 references, so people have been interested in this for a while! Also, PX in digital games has produced many questionnaires for similar constructs.
The main body of the paper envisions which aspects of the TTRPG PX can be tracked (and how they can be impacted by GMs & game rules). We identify 1/ Cognitive Challenge (strategy, cognitive dissonance, memory) 2/ Immersion 3/ Agency 4/ Attachment 5/ Group Dynamics 6/ Refereeing.
Why should TTRPG stakeholders care about validated instruments to measure PX? GMs can use them to solicit feedback & track PX across sessions or when trying new rules. TTRPG designers can collect more informed & comparable playtest data. Same for event organizers! (img: PaizoCon)