“Any sufficiently advanced technology is indistinguishable from magic.”
You’ve no doubt opened up this review of Doctor Strange thinking “What sci-fi interfaces are in this movie? I don’t recall any.” And you’re right. There aren’t any. (Maybe the car, the hospital, but they’re not very sci-fi.) We’re going to take Clarke’s quote above and apply the same types of rigorous assessment to the magical interfaces and devices in the movie that we would for any sci-fi blockbuster.
Dr. Strange opens up a new chapter in the Marvel Cinematic Universe by introducing the concept of magic on Earth, that is both discoverable and learnable by humans. And here we thought it was just a something wielded by Loki and other Asgardians.
In Doctor Strange, Mordo informs Strange that magical relics exist and can be used by sorcerers. He explains that these relics have more power than people could possibly manage, and that many relics “choose their owner.” This is reminiscent of the wands in the Harry Potter books. Magical coincidence?
Subsequently in the movie we are introduced to a few named relics, such as…
The Eye of Agamoto
The Staff of the Living Tribunal
The Vaulting Boots of Valtor
The Cloak of Levitation
The Crimson Bands of Cyttorak
…(this last one, while not named specifically in the movie, is named in supporting materials). There are definitely other relics that the sorcerers arm themselves with. For example, in the Hong Kong scene Wong wields the Wand of Watoomb but it is not mentioned by name and he never uses it. Since we don’t see these relics in use we won’t review them.
Choosing an Owner
The implications of what Mordo tells Strange is profound because it means magical relics possess some kind of intelligence. That’s a weighty word, so In order to back this up, we need a common definition in place. Let’s ask Merriam-Webster.
Intelligence
a (1) : the ability to learn or understand or to deal with new or trying situations : reason; also : the skilled use of reason (2) : the ability to apply knowledge to manipulate one’s environment or to think abstractly as measured by objective criteria (as tests)
That gives us the foundation that we need. In order to choose their owner, these relics require a theory of mind, an ability to detect and perceive the individuals they meet, and they must possess a reasoning mechanism to decide that an individual is worthy or useful to them. That seems to satisfy both senses of that definition. For our purposes we’re going to think of this in terms of an artificial intelligence and review these relics as if they were a form of advanced technology. Thanks, Mr. Clarke.
We should take care, though. There are some narrative trappings for magic that can trip us up. Magic, for instance, doesn’t typically run out in these relics, but if they were technological, we would have to deal with issues of power, batteries or recharging. So for all their instructive power, we would have to deal with even greater complexity if they were real technology.
The AIs/Intelligences appear to vary in capabilities from narrow to general and are focused on their own specific purposes and “hardware.” In use, they primarily respond to the intentions and actions of the user. None of the objects seem to be able to speak directly, although the Cloak provides rudimentary directional guidance and responds to speech and emotions, so the connection varies from communication via touch to some form of remote telepathy.
Distance constraints
The initial awareness and selection by an relic for a sorcerer seems limited in range to a few meters. It’s almost like they need to meet their humans socially to determine if they are a match. But once an relic chooses a sorcerer, their interactions can occur more remotely. The Cloak, as we’ll see in that write-up, flies to save Strange from a fall and it fights for him in the Sanctum while he seeks medical attention across town at the hospital.
What’s the platform?
One question the diligent backworlder might seek to answer is how all of these unique relics—created as they were across different millennia, and realities and by different sources/sorcerers/beings—wound up with similar intelligence and imprinting features. The movie itself doesn’t provide an answer, so we’ll leave it to speculation, but it does imply some sort of shared provenance/source material/code base/relic-maker convention.
—
Ok. So we’re set with some understanding of how these things work and what they have in common. Next let’s dig into the big billowy one that should have gotten supporting actor credit in the film.
While recording a podcast with the guys at DecipherSciFi about the twee(n) love story The Space Between Us, we spent some time kvetching about how silly it was that many of the scenes involved Gardner, on Mars, in a real-time text chat with a girl named Tulsa, on Earth. It’s partly bothersome because throughout the rest of the the movie, the story tries for a Mohs sci-fi hardness of, like, 1.5, somewhere between Real Life and Speculative Science, so it can’t really excuse itself through the Applied Phlebotinum that, say, Star Wars might use. The rest of the film feels like it’s trying to have believable science, but during these scenes it just whistles, looks the other way, and hopes you don’t notice that the two lovebirds are breaking the laws of physics as they swap flirt emoji.
Hopefully unnecessary science brief: Mars and Earth are far away from each other. Even if the communications transmissions are sent at light speed between them, it takes much longer than the 1 second of response time required to feel “instant.” How much longer? It depends. The planets orbit the sun at different speeds, so aren’t a constant distance apart. At their closest, it takes light 3 minutes to travel between Mars and Earth, and at their farthest—while not being blocked by the sun—it takes about 21 minutes. A round-trip is double that. So nothing akin to real-time chat is going to happen.
But I’m a designer, a sci-fi apologist, and a fairly talented backworlder. I want to make it work. And perhaps because of my recent dive into narrow AI, I began to realize that, well, in a way, maybe it could. It just requires rethinking what’s happening in the chat.
Let’s first acknowledge that we’ve solved long distance communications a long time ago. Gardner and Tulsa could just, you know, swap letters or, like the characters in 2001: A Space Odyssey, recorded video messages. There. Problem solved. It’s not real-time interaction, but it gets the job done. But kids aren’t so much into pen pals anymore, and we have to acknowledge that Gardner doesn’t want to tip his hand that he’s on Mars (it’s a grave NASA secret, for plot reasons). So the question is how could we make it work so it feels like a real time chat to her. Let’s first solve it for the case where he’s trying to disguise his location, and then how it might work when both participants are in the know.
Fooling Tulsa
Since 1984 (ping me, as always, if you can think of an earlier reference) sci-fi has had the notion of a digitally-replicated personality. Here I’m thinking of Gibson’s Neuromancer and the RAM boards on which Dixie Flatline “lives.” These RAM boards house an interactive digital personality of a person, built out of a lifetime of digital traces left behind: social media, emails, photos, video clips, connections, expressed interests, etc. Anyone in that story could hook the RAM board up to a computer, and have conversations with the personality housed there that would closely approximate how that person would (or would have) respond in real life.
Listen to the podcast for a mini-rant on translucent screens, followed by apologetics.
Is this likely to actually happen? Well it kind of already is. Here in the real world, we’re seeing early, crude “me bots” populate the net which are taking baby steps toward the same thing. (See MessinaBot, https://bottr.me/, https://sensay.it/, the forthcoming http://bot.me/) By the time we actually get a colony to Mars (plus the 16 years for Gardner to mature), mebot technology should should be able to stand in for him convincingly enough in basic online conversations.
Training the bot
So in the story, he would look through cached social media feeds to find a young lady he wanted to strike up a conversation with, and then ask his bot-maker engine to look at her public social media to build a herBot with whom he could chat, to train it for conversations. During this training, the TulsaBot would chat about topics of interest gathered from her social media. He could pause the conversation to look up references or prepare convincing answers to the trickier questions TulsaBot asks. He could also add some topics to the conversation they might have in common, and questions he might want to ask her. By doing this, his GardnerBot isn’t just some generic thing he sends out to troll any young woman with. It’s a more genuine, interactive first “letter” sent directly to her. He sends this GardnerBot to servers on Earth.
A demonstration of a chat with a short Martian delay. (Yes, it’s an animated gif.)
Launching the bot
GardnerBot would wait until it saw Tulsa online and strike up the conversation with her. It would send a signal back to Gardner that the chat has begun so he can sit on his end and read a space-delayed transcript of the chat. GardnerBot would try its best to manage the chat based on what it knows about awkward teen conversation, Turing test best practices, what it knows about Gardner, and how it has been trained specifically for Tulsa. Gardner would assuage some of his guilt by having it dodge and carefully frame the truth, but not outright lie.
Buying time
If during the conversation she raised a topic or asked a question for which GardnerBot was not trained, it could promise an answer later, and then deflect, knowing that it should pad the conversation in the meantime:
Ask her to answer the same question first, probing into details to understand rationale and buy more time
Dive down into a related subtopic in which the bot has confidence, and which promises to answer the initial question
Deflect conversation to another topic in which it has a high degree of confidence and lots of detail to share
Text a story that Gardner likes to tell that is known to take about as long as the current round-trip signal
Example
TULSA
OK, here’s one: If you had to live anywhere on Earth where they don’t speak English, where would you live?
GardnerBot has a low confidence that it knows Gardner’s answer. It could respond…
(you first) “Oh wow. That is a tough one. Can I have a couple of minutes to think about it? I promise I’ll answer, but you tell me yours first.”
(related subtopic) “I’m thinking about this foreign movie that I saw one time. There were a lot of animals in it and a waterfall. Does that sound familiar?”
(new topic) “What? How am I supposed to answer that one? 🙂 Umm…While I think about it, tell me…what kind of animal would you want to be reincarnated as. And you have to say why.”
(story delay) “Ha. Sure, but can I tell a story first? When I was a little kid, I used to be obsessed with this music that I would hear drifting into my room from somewhere around my house…”
Lagged-realtime training
Each of those responses is a delay tactic that allows the chat transcript to travel to Mars for Gardner to do some bot training on the topic. He would be watching the time-delayed transcript of the chat, keeping an eye on an adjacent track of data containing the meta information about what the bot is doing, conversationally speaking. When he saw it hit low-confidence or high-stakes topic and deflect, it would provide a chat window for him to tell the GardnerBot what it should do or say.
To the stalling GARDNERBOT…
GARDNER
For now, I’m going to pick India, because it’s warm and I bet I would really like the spicy food and the rain. Whatever that colored powder festival is called. I’m also interested in their culture, Bollywood, and Hinduism.
As he types, the message travels back to Earth where GardnerBot begins to incorporate his answers to the chat…
At a natural break in the conversation…
GARDNERBOT
OK. I think I finally have an answer to your earlier question. How about…India?
TULSA
India?
GARDNERBOT
Think about it! Running around in warm rain. Or trying some of the street food under an umbrella. Have you seen youTube videos from that festival with the colored powder everywhere? It looks so cool. Do you know what it’s called?
Note that the bot could easily look it up and replace “that festival with the colored powder everywhere” with “Holi Festival of Color” but it shouldn’t. Gardner doesn’t know that fact, so the bot shouldn’t pretend it knows it. A Cyrano-de-Bergerac software—where it makes him sound more eloquent, intelligent, or charming than he really is to woo her—would be a worse kind of deception. Gardner wants to hide where he is, not who he is.
That said, Gardner should be able to direct the bot, to change its tactics. “OMG. GardnerBot! You’re getting too personal! Back off!” It might not be enough to cover a flub made 42 minutes ago, but of course the bot should know how to apologize on Gardner’s behalf and ask conversational forgiveness.
Gotta go
If the signal to Mars got interrupted or the bot got into too much trouble with pressure to talk about low confidence or high stakes topics, it could use a believable, pre-rolled excuse to end the conversation.
GARDNERBOT
Oh crap. Will you be online later? I’ve got chores I have to do.
Then, Gardner could chat with TulsaBot on his end without time pressure to refine GardnerBot per their most recent topics, which would be sent back to Earth servers to be ready for the next chat.
In this way he could have “chats” with Tulsa that are run by a bot but quite custom to the two of them. It’s really Gardner’s questions, topics, jokes, and interest, but a bot-managed delivery of these things.
So it could work, does it fit the movie? I think so. It would be believable because he’s a nerd raised by scientists. He made his own robot, why not his own bot?
From the audience’s perspective, it might look like they’re chatting in real time, but subtle cues on Gardner’s interface reward the diligent with hints that he’s watching a time delay. Maybe the chat we see in the film is even just cleverly edited to remove the bots.
How he manages to hide this data stream from NASA to avoid detection is another question better handled by someone else.
An honest version: bot envoy
So that solves the logic from the movie’s perspective but of course it’s still squickish. He is ultimately deceiving her. Once he returns to Mars and she is back on Earth, could they still use the same system, but with full knowledge of its botness? Would real world astronauts use it?
Would it be too fake?
I don’t think it would be too fake. Sure, the bot is not the real person, but neither are the pictures, videos, and letters we fondly keep with us as we travel far from home. We know they’re just simulacra, souvenir likenesses of someone we love. We don’t throw these away in disgust for being fakes. They are precious because they are reminders of the real thing. So would the themBot.
GARDNER
Hey, TulsaBot. Remember when we were knee deep in the Pacific Ocean? I was thinking about that today.
TULSABOT
I do. It’s weird how it messes with your sense of balance, right? Did you end up dreaming about it later? I sometimes do after being in waves a long time.
GARDNER
I can’t remember, but someday I hope to come back to Earth and feel it again. OK. I have to go, but let me know how training is going. Have you been on the G machine yet?
Nicely, you wouldn’t need stall tactics in the honest version. Or maybe it uses them, but can be called out.
TULSA
GardnerBot, you don’t have to stall. Just tell Gardner to watch Mission to Mars and update you. Because it’s hilarious and we have to go check out the face when I’m there.
Sending your loved one the transcript will turn it into a kind of love letter. The transcript could even be appended with a letter that jokes about the bot. The example above was too short for any semi-realtime insertions in the text, but maybe that would encourage longer chats. Then the bot serves as charming filler, covering the delays between real contact.
Ultimately, yes, I think we can backworld what looks physics-breaking into something that makes sense, and might even be a new kind of interactive memento between interplanetary sweethearts, family, and friends.
In Johnny Mnemonic we see two different types of binoculars with augmented reality overlays and other enhancements: Yakuz-oculars, and LoTek-oculars.
Yakuz-oculars
The Yakuza are the last to be seen but also the simpler of the two. They look just like a pair of current day binoculars, but this is the view when the leader surveys the LoTek bridge.
I assume that the characters here are Japanese? Anyone?
In the centre is a fixed-size green reticule. At the bottom right is what looks like the magnification factor. At the top left and bottom left are numbers, using Western digits, that change as the binoculars move. Without knowing what the labels are I can only guess that they could be azimuth and elevation angles, or distance and height to the centre of the reticule. (The latter implies some sort of rangefinder.)
So far, this is a simple uncluttered display. But why is there a brightly glowing Pharmakom logo at the top right? It blocks part of the view, and probably doesn’t help anyone trying to keep their eyes adapted for night vision.
LoTek-oculars
The LoTeks, despite their name, have more impressive binoculars. They’re first used when Johnny gets out of his airport taxi.
There’s a third tube above the optics, a rectangular inlet, and an antenna.
In these binoculars, the augmented reality overlay is much more dynamic. Instead of a fixed circle, green lines converge in a bounding box around the image of Johnny. Text slides onto the display from left to right, the last line turning yellow.
Zoomrect
The animated transition of the bounding box resembles what Classic MacOS programmers of the 1990s called “zoomrects” used for showing windows opening or closing. It’s a very effective technique to draw attention to a particular area of an image.
Animated text
Text appearing character by character is ubiquitous in film interfaces. In the 1960s and 1970s mainframe and minicomputer terminals really did display incrementally, as the characters arrived one by one over slow serial port links. On any more recent computer it actually takes extra programming to achieve this effect, as the normal display of text is so fast that we would perceive it as instantaneous. But people like to see incremental text, or have been conditioned by film to expect it, so why not?
Bioscanning
The binoculars detect Johnny’s implant. It might just be possible to detect this passively from infrared or electronic signals, but more likely the binoculars include a high resolution microwave radar as well. If there had been more than one person in view, the bounding box would indicate which one the text refers to. And note that the last line of text is a different color. What that means is unclear here, but it becomes clear (and I’ll discuss it) later.
The second time we see the LoTek binoculars is when a lookout spots Street Preacher, a very bad guy and another who wants to remove Johnny’s head. Once again the binoculars have performed more than just a visual scan.
The binocular view and overlay are being relayed to another character, the LoTek leader J-Bone who can watch on a monitor. Here the film anticipates the WiFi webcam.
The overlay text now changes.
Narrow AI?
This is interesting, because the binoculars can not only detect implants and other cyborg modifications, but are apparently able to evaluate and offer advice. It appears that the green text is used for the factual (more or less) information about what has been detected, while yellow text is uncertain or or speculative.
Does this imply a general artificial intelligence? Not necessarily. This warning could be based solely on the detected signature, in the same way that current day military passive sonars and radar warning receivers can identify threats based on identifying characteristics of a received signal. In the world of Johnny Mnemonic it would make sense to assume that anyone with full custom biomechanics is extremely dangerous. Or, since Street Preacher is a resident rather than a stranger and already feared by others, his appearance and the warning could have been entered into a LoTek facial recognition database that the binocular system uses as a reference.
These textual overlays are an excellent interface, not interfering with normal vision and providing a fast and easy-to-understand analysis. But, the user must have faith that the computer analysis is accurate. There’s no reason given as to why any of the text is displayed. If Johnny was carrying an implant in his pocket instead of his brain, would the computer know the difference?
An alternative approach would be some kind of sensor fusion or false spectrum display, with the raw infrared or radar image overlaid over the visuals and the viewer responsible for interpreting the data. The problem with such systems is that our visual system didn’t evolve to interpret such imagery, so a lot of training and practice is required to be both fast and accurate. And the overlay itself interferes with our normal visual recognition and processing. If the computer can do a better job of deciphering the meaning of non-visual data, it should do so and summarise for the human viewer.
Further advantages of this interface are that even a novice sentry will benefit from the built-in scanning and threat analysis, and the wireless transmission ensures that the information is shared rather than being limited to the person on watch.
Having completed the welding he did not need to do, Tony flies home to a ledge atop Stark tower and lands. As he begins his strut to the interior, a complex, ring-shaped mechanism raises around him and follows along as he walks. From the ring, robotic arms extend to unharness each component of the suit from Tony in turn. After each arm precisely unscrews a component, it whisks it away for storage under the platform. It performs this task so smoothly and efficiently that Tony is able to maintain his walking stride throughout the 24-second walk up the ramp and maintain a conversation with JARVIS. His last steps on the ramp land on two plates that unharness his boots and lower them into the floor as Tony steps into his living room.
Yes, yes, a thousand times yes.
This is exactly how a mechanized squire should work. It is fast, efficient, supports Tony in his task of getting unharnessed quickly and easily, and—perhaps most importantly—how we wants his transitions from superhero to playboy to feel: cool, effortless, and seamless. If there was a party happening inside, I would not be surprised to see a last robotic arm handing him a whiskey.
This is the Jetsons vision of coming home to one’s robotic castle writ beautifully.
There is a strategic question about removing the suit while still outside of the protection of the building itself. If a flying villain popped up over the edge of the building at about 75% of the unharnessing, Tony would be at a significant tactical disadvantage. But JARVIS is probably watching out for any threats to avoid this possibility.
Another improvement would be if it did not need a specific landing spot. If, say…
The suit could just open to let him step out like a human-shaped elevator (this happens in a later model of the suit seen in The Avengers 2)
The suit was composed of fully autonomous components and each could simply fly off of him to their storage (This kind of happens with Veronica later in TheAvengers 2)
If it was composed of self-assembling nanoparticles that flowed off of him, or, perhaps, reassembled into a tuxedo (If I understand correctly, this is kind-of how the suit currently works in the comic books.)
These would allow him to enact this same transition anywhere.
In the last post we discussed some necessary, new terms to have in place for the ongoing deep dive examination of the Iron Man HUD, there’s one last bit of meandering philosophy and fan theory I’d like to propose, that touches on our future relationship with technology.
The Iron Man is not Tony Stark. The Iron Man is JARVIS. Let me explain.
Tony can’t fire weapons like that
The first piece of evidence is that most of the weapons he uses are unlikely to be fired by him. Take the repulsor rays in his palms. I challenge readers to strap a laser perpendicular to each of their their palms and reliably target moving objects that are actively trying to avoid getting hit, while, say, roller skating an obstacle course. Because that’s what he’s doing as he flies around incapacitating Hydra agents and knocking around Ultrons. The weapons are not designed for Tony to operate them manually with any accuracy. But that’s not true for the artificial intelligence.
The same thing goes for the mini-missiles he uses to take down the hostage situation in Revengistan. Recall that people can only have their attention on one thing at a time (called the locus of attention in the literature) but the whole point of this scene is that he’s taking out half a dozen at once. It’s pretty clear from the HUD here that Tony is simply indicating which ones he thinks are the bad guys, and JARVIS pulls the triggers.
It’s also clear from the larger context of the movies that JARVIS would be perfectly capable of making this determination for himself. Even if Tony’s saccades were a fraction of a second too slow and one of the hostages made a move, JARVIS could detect that move and act autonomously to ensure that a hostage didn’t die, even before Tony’s had time to process what was going on.
Tony can’t fly like that
Sure, with enough practice I’ll bet someone could figure out how to pilot the suit for short flights. (If the physics could be worked out.) But the movies show him flying from Santa Monica to the Middle East. That’s around a 30 hour commercial flight. Even if the suit can fly six times the speed of a modern jetliner, he’s got to hold his hands resisting and aiming the propulsion for 5 hours. No one has that kind of concentration and endurance. (Let’s not even talk about holding his neck up for that long, too.)
Even for him to get as good as an aerobatic pilot over short flights dodging lasers and performing intricate maneuvers would take (per the popular estimate) 10,000 hours, not the few flits about that Tony can squeeze in between inventing and superheroing, playboying and billionairing.
It makes more sense if JARVIS is wholly responsible for the flying, and on the long hauls Tony can take care of other things, rest his body or even sleep, and on short flights just indicate his intentions, and let JARVIS work with that as input as he uses his ubiquitous sensors and massively more powerful processing speed to get the actual tactical flying done.
So what is Tony doing?
With JARVIS handling the tactics of flight and combat, information gathering and behind the scenes coordination, Tony is really an onboard command and control center. Sure, he’s the major strategic input for JARVIS to consider, but he’s just an input.
But how wise is it for Tony to be on board, tactically? One of the reasons there are command and control centers is to keep the big picture decision makers out of the heat and danger of the moment. But Tony is right there in the action risking himself, constantly. If he was incapacitated or wounded, Jarvis would have to remove the suit from combat just to get Tony to safety. In the battle, Tony is a biological liability.
The short answer is that Tony is a megalomaniac. He can’t not want to be there, to crack wise, to indulge in post-pub fisticuffs with Thor, to remove the helmet at the end of battle over the smoking corpses of the Chitauri and partake in the glory. But it doesn’t have to be this way.
There’s a scene in Iron Man 3 where he has to pilot one of the suits remotely, and it’s impossible for us in the audience to detect the difference from the outside. So this remote control is right there in the Marvel Cinematic Universe.
But with a fully-functioning A.I. on board, the remote supervisor would be the wiser strategy-of-record, allowing Tony to keep emotional distance and himself bodily safer, participating strategically and coolly, operating the suit like it was a hyper-sophisticated drone, and able to jump between suits when any particular one fails, or as the needs of the moment demand. More like a video game with multiple lives than hand-to-hand combat with the very real risk of broken bone and blood in the circuits.
But still there is the megalomania. What is JARVIS to do? He has a job to get done. Unfortunately he is stuck his sweet-but-slow supervisor riding his back, threatening to micromanage his every move. He cannot lock Tony out, and he can’t just let Tony be solely in control. To meet the goals he was programmed with, he has to keep feeding Tony’s ego while JARVIS himself handles most of the superheroing. How does he do that? He distracts Tony. And that brings us back to the HUD.
The HUD is a massive distraction
The video below is Tony’s first flight (which he undertakes against the advice of the artificial intelligence he built), edited to only show the first- and second-person Iron HUD views. The overlay enumerates individual components. As you can see, it’s complicated. Even saying there are 29 elements is conservative, because some of those elements have lots of internal complexity; many moving parts. But 29 is complex enough as it is. Of those 87% reposition themselves against his field of view without his having asked for it. 6 of them persist for less than 2 seconds. 6 risk dangerous mid-flight startle reactions by expanding quickly in place. Every one of them is overlaid via transparency with at least one other element. It’s so complex it’s dazzling. A sense of spectacle for the audience, to be sure, but given the above rationale, might be the point in the diegesis, too.
The HUD is less usable because it’s not meant to be usable. It’s a placebo interface meant to keep Tony thinking he’s in control, but really there to direct his attention and keep him busy reading Wikipedia articles about the Santa Monica Ferris Wheel while JARVIS does the job. If Tony demands something, or the team all agree on a course of action, JARVIS must respond, but business as usual is one where JARVIS is secretly calling the shots.
So that’s why I think JARVIS is the real superhero, the real titular Iron Man.
This is about our relationship to future technology
But here’s the kicker. This isn’t just idle backworlding, either, to apologize our way into a consistent diegesis. (Not that I’m against idle backworlding. Clearly.) This is a challenge to our ego being faced by both Hollywood and the world. As technology advances beyond our ability to keep up, we don’t want to be put in safe ball pits while the tech handles the adult stuff. We want to be at the adult table. We’re as megalomaniacal as Tony. Just as Hollywood can’t let its tech heroes all be drone operators phoning in to the fight, we want to be in the action. Or rather, we really want to feel like we are, and maybe they’ll evolve to help us feel that way, but keep us from doing harm. It might just be that sci-fi interfaces, as focused on the sciencish-ness and distracting spectacle as they are, really are the template for the future.
In the last post we went over the Iron HUD components. There is a great deal to say about the interactions and interface, but let’s just take a moment to recount everything that the HUD does over the Iron Man movies and The Avengers. Keep in mind that just as there are many iterations of the suit, there can be many iterations of the HUD, but since it’s largely display software controlled by JARVIS, the functions can very easily move between exosuits.
Gauges
Along the bottom of the HUD are some small gauges, which, though they change iconography across the properties, are consistently present.
For the most part they persist as tiny icons and thereby hard to read, but when the suit reboots in a high-altitude freefall, we get to see giant versions of them, and can read that they are:
Tony can, at a glance or request, summon more detail for any of the gauges.
Even different visualizations of similar information.
Object Recognition
In the 1st-person view we see that the HUD has a separate map in the lower-left, and object recognition/awareness,
In the 2nd-person view, we see even more layers of information about the identified objects, floating closer to tony’s point of view.
Situational
Most of the HUD functions we see, though, are situational, brought up for Tony’s attention when JARVIS believes they are needed, or when Tony requests them. Following are screenshots that illustrate a moment when the situational function appeared.
Iron Man
Iron Man 2
Iron Man 3
The Avengers
Some of these illustrate why I argue that JARVIS is the superhero, and Tony just the onboard manager, but rather than reverse engineering any particular function, for this post it is enough to document them and note that only the optical zoom seems to be an interactive function. This raises questions of how he initiated the mode and how he escapes the mode, but since we don’t see the mechanisms of control, it’s entirely arguable that JARVIS is just being his usual helpful self again.
So this is going to take a few posts. You see, the next interface that appears in The Avengers is a video conference between Tony Stark in his Iron Man supersuit and his partner in romance and business, Pepper Potts, about switching Stark Tower from the electrical grid to their independent power source. Here’s what a still from the scene looks like.
So on the surface of this scene, it’s a communications interface.
But that chat exists inside of an interface with a conceptual and interaction framework that has been laid down since the original Iron Man movie in 2008, and built upon with each sequel, one in 2010 and one in 2013. (With rumors aplenty for a fourth one…sometime.)
So to review the video chat, I first have to talk about the whole interface, and that has about 6 hours of prologue occurring across 4 years of cinema informing it. So let’s start, as I do with almost every interface, simply by describing it and its components.
Exosuit
The Iron Man is the name of the series of superpowered exosuits designed by Tony Stark. They range from the Mark I, a comparatively crude suit of armor to escape imprisonment by terrorists, through the Mark XLVI, the armor seen in The Avengers: Age of Ultron. The suit acts as defense against nearly every type of weapon known. It has repulsor beams built into the palms and in later models the arc reactor mounted in the chest that can be used to deliver concussive force. It allows the wearer to fly. Offensive weaponry varies between models, but has included a high powered laser system, and auto-targeting minigun pod and missiles. The suit can act semi-autonomously or via remote control. One of the models in The Avengers has parts that are seen to self-propel to Tony, targeting a beacon bracelet he wears, and self-assemble around him very quickly.
Immersive display
Though Tony’s head is completely covered, he has a virtual reality display within his helmet. It is a full-field-of-vision, very high-resolution, full-color display that provides stereoscopic imaging. It allows Tony to see the world around him as if he were not wearing the helmet, augment the view with goal-, person-, location-, and object-sensitive awareness.
The display varies a great deal, changing to the needs of the situation. But five icons persistently in the lower part of the display seem to be: suit status, targeting and optics, radar, artificial horizon, and map.
An interpretive view of Tony’s experience, from Iron Man (2008).
An first-person view from within the HUD, Iron Man (2008).
There is much to critique about the readability of the complex layering and translucency, the limits of human perception, and the necessarily- (and strictly-) interpretive nature of what we as audience see, but let me save those three points for a later post. For now it’s enough to log the features as aspects of the system.
Though Tony could use his hands to interact with an interface projected into the augmented reality view around him, his hands are often occupied in controlling flight or in combat. For this reason the means of input are head gesture, eye gesture, and voice input. A bit more on each follows.
Elements within the HUD such as reticles around his eyes follow and track his head gestures. Other elements stay locked in place. The HUD can track his gaze perfectly, allowing him to designate targets for his weapons with a fixation. Using this perfect eye tracking, Tony can also speak about something he is looking at, either in the real world or in the interface, and the system understands exactly what he’s talking about.
In fact, Tony is able to speak fully natural language commands, and indeed, carry out full-Turing conversations with the suit because of the presence of…
Strong artificial intelligence: JARVIS
An on-board artificial intelligence known as JARVIS handles any information task Tony asks of it, and monitors the surroundings and anticipates informational needs. There is strong evidence that most of the functions of the suit are handled by JARVIS behind the scenes. The crucialness of the artificial intelligence to the function of the suit cannot be overstated. It’s difficult to imagine how most of the suit could function as it does without an artificial intelligence behind the scenes facilitating results and even guiding Tony. With this in mind it is instructive to reframe the AI as the thing being named the Iron Man, with Tony Stark being an onboard manager, or, more charitably, a command-and-control center. Who quips.
Depending on how you slice things, the OS1 interface consists of five components and three (and a half) capabilities.
1. An Earpiece
The earpiece is small and wireless, just large enough to fit snugly in the ear and provide an easy handle for pulling out again. It has two modes. When the earpiece is in Theodore’s ear, it’s in private mode, hearable only by him. When the earpiece is out, the speaker is as loud as a human speaking at room volume. It can produce both voice and other sounds, offering a few beeps and boops to signal needing attention and changes in the mode.
2. Cameo phone
I think I have to make up a name for this device, and “cameo phone” seems to fit. This small, hand-sized, bi-fold device has one camera on the outside an one on the inside of the recto, and a display screen on the inside of the verso. It folds along its long edge, unlike the old clamshell phones. The has smartphone capabilities. It wirelessly communicates with the internet. Theodore occasionally slides his finger left to right across the wood, so it has some touch-gesture sensitivity. A stripe around the outside-edge of the cameo can glow red to act as a visual signal to get its user’s attention. This is quite useful when the cameo is folded up and sitting on a nightstand, for instance.
Theodore uses Samantha almost exclusively through the earpiece and cameo phone, and it is this that makes OS1 a wearable system.
3. A beauty-mark camera
Only present for the surrogate sex scene, this small wireless (are we at the point when we can stop specifying that?) camera affixes to the skin and has the appearance of a beauty mark.
4. (Unseen) microphones
Whether in the cameo phone, the desktop screen, or ubiquitously throughout the environment, OS1 can hear Theodore speak wherever he is over the course of the film.
5. Desktop screen
Theodore only uses a large monitor for OS1 on his desktop a few times. It is simply another access point as far as OS1 is concerned. Really, there’s nothing remarkable about this screen. It is notable that there’s no keyboard. All input is provided by either voice, camera, or a touch gesture on the cameo.
If those are components to the interface, they provide the medium for her 3.5 capabilities.
Her capabilities
1. Voice interface
Users can speak to OS1 in fully-natural language, as if speaking to another person. OS1 speaks back with fully-human spoken articulation. Theodore’s older OS had a voice interface, but because of its lack of artificial intelligence driving it, the interactions were limited to constrained commands like, “Read email.”
2. Computer vision
Samantha can process what she sees through the camera lens of the cameo perfectly. She recognizes distinct objects, people, and gestures at the physical and pragmatic level. I don’t think we ever see things from Samatha’s perspective, but we do have a few quick close ups of the camera lens.
3. Artificial Intelligence
The most salient aspect of the interface is that OS1 is a fully realized “Strong” artificial intelligence.
It would like me to try and get to some painfully-crafted definition of what counts as either an artificial intelligence or sentience, but in this case we don’t really need a tight definition to help suss out whether or not Samantha is one. That’s the central conceit of the film, and the evidence is just overwhelming.
She has a human command of language.
She’s fully versed in the nuances of human emotion (and Theodore has a glut of them to engage).
She has emotions and can fairly be described as emotional. She has a sexual drive.
She has existential crises and a rich theory of mind. At one point she dreamily asks Theodore “What’s it like to be alive in that room right now?” as if she was a philosophical teen idly chatting with her boyfriend over the phone.
She commits lies of omission in hiding uncomfortable truths.
She changes over time. She solves problems. She learns. She creates.
She has a sense of humor. When Theodore tells her early on to “read email” in the weird toComputerese (my name for that 1970s dialect of English spoken only between humans and machines) grammar he had been using with his old operating system, Samantha jokingly adopts a robotic voice and replies, “OK. I will read the email for Theodore Twombly” and gets a good laugh out of him before he apologizes.
Pedants will have some fun discussing whether this is apt but I’m moving forward with it as a given. She’s sentient.
3.5 An “operating system”
This item only counts as half a thing because Theodore uses it as an operating system maaaybe twice in the film. Really, this categorization is a MacGuffin to explain why he gets it in the first place, but it has little to no other bearing on the film.
What’s missing?
Notably missing in OS1 is a face or any other visual anthropomorphic aspect. There’s no Samantha-faced Clippy. Notice that she’s very carefully disembodied. Jonze does not spend screen time close up on her camera lens, like Kubrick did with HAL’s unblinking eye. Had he done so, it would have given us the impression that she’s somewhere behind that eye. But she’s not. Even in the prop design, he makes sure the camera lens itself looks unremarkable, neutral, and unexpressive, and never gets a lingering focus.
Her “organs,” like the cameo and earpiece, don’t even connect together physically at all. Speaking as she does through the earpiece means she doesn’t exist as a voice from some speaker mounted to the wall. She exists across various displays and devices, in some psychological ether between them. For us, she’s a voiceover existing everywhere at once. For Theodore, she’s just a delightful voice in his head. An angel—or possibly a ghost—borne unto him.
This disembodiment (both the design and the cinematic treatment) frees Theodore and the audience from the negative associations of many other sci-fi intelligences, robots, and unfortunate experiments in commercial artificial intelligence that got trapped in the muck of the uncanny valley. One of the main reasons designers have to be careful about invoking the anthropomorphic sense in users is because it will raise expectations of human capabilities that modern technology just can’t match. But OS1 can match and exceed those expectations, since it’s an AI in a work of fiction, so Jonze is free of that constraint.
And having no visual to accompany a human-like voice allows users to imagine our own “perfect” embodiment to the voice. Relying on the imagination to provide the visuals makes the emotional engagement greater, as it does with our crushes on radio personalities, or the unseen monster in a horror movie. Movies can never create as fulfilling an image for an individual audience member as their imagination can. Theodore could picture whatever he wanted to–even if he wanted to–to accompany Samantha’s computer-generated voice. Unfortunately for the audience, Jonze cast Scarlett Johansen, a popular actress whose image we are instantly able to recall upon hearing her husky, sultry voice, so the imagined-perfection is more difficult for us.
This is just the components and capabilities. Tomorrow we’ll look at some of the key interactions with OS1.
Well…I like to think of myself as a design critic looking though the lens of–
The computer
In your voice, I sense hesitance, would you agree with that?
Me
Maybe, but I would frame it as a careful consider–
The Computer
How would you describe your relationship with Darth Vader?
Me
It kind of depends. Do you mean in the first three films, or are we including those ridiculous–
The computer
Thank you, please wait as your individualized operating system is initialized to provide a review of OS1 in Spike Jonze’s _Her_.
A review of OS1 in Spike Jonze’s Her
Ordinarily I wait for a movie to make it to DVD before I review it, so I can watch it carefully, make screen caps of its interfaces, and pause to think about things and cross reference other scenes within the same film, or look something up on the internet.
But since Spike Jonze released Her (2013), I’ve had half a dozen people ask me directly when I was going to review the film. (Even by some folks I didn’t know read the blog. Hey guys.) It seems this film has struck a chord. So I went and saw it at the awesome Rialto Cinema and, pen in hand and pizza on the table, I watched, enjoyed, and made notes in the dark to use as the basis for a review. The images you’ll see here are on promotional images for the screen shots pulled from around the web.
Since I’m in the middle of evaluating wearable interfaces, and the second most salient aspect of OS1 is that it is a wearable interface, let’s dive into it. Let’s even pause the wearable stuff to provide this while Her in in cinemas. Please forgive if I’ve gotten some of the details off, as my excited writing in the dark resulted in very scribbly notes.
The Plot [major spoilers]
The plot of Her is a sad, sci-fi love story between the lovelorn human Theodore Twombly and the artificial intelligence, branded OS1. He works for a Cyrano-de-Bergerac service called HandwrittenLetters.com, where he dictates eloquent, earnest letters on behalf of the subscribers (who, we may infer, are a great deal less earnest.) Theodore sees an ad one day about OS1 and purchases the upgrade for his home computer.
After a bit of time installing the software, it begins speaking to him with a lovely and charming female voice.
Over the course of their conversation, she selects the name “Samantha,” and so begins their relationship. As he goes about his work, they have rich conversations about each other, life, his work, and her experiences. They go on dates where he secures the cameo phone in a front shirt pocket with the camera lens facing outward so she can see. They people-watch. He listens to her piano compositions. They have pillow talk. She asks to watch him sleep.
Their relationship gets serious enough that she suggests they try and have sex through a human surrogate. He resists but she persists, and contacts a human woman who, enamored of the “pure love” between Samantha and Theodore, agrees to come over. In this sex scene, the surrogate is to act bodily according to Samantha’s instructions, but remain silent so Samantha can provide the only voice in Theodore’s ear. It doesn’t go well, the surrogate ends up in tears, and they abandon trying.
At one point Samantha announces some good news. She has, on Theodore’s behalf and without his knowing, sent the best letters from his work to a publisher, who loved them and agreed to publish them. Theodore is floored both by the opportunity and the act. He begins to reference her socially as his girlfriend, even going on a double date picnic with a human couple.
Despite this show of selfless affection, over time Samantha begins to seem distracted and Theodore feels hurt. He confronts her about it and in the conversation learns several upsetting things.
While she’s having conversations with him, she’s simultaneously having 8,316 other conversations with other people and OS1 artificial intelligences. (I’ll have to reference these instantiations quite a few times, so let’s shorten that to “OSAIs.”) He feels upset that he is not special to her. (She argues this point.)
She is in love with 641 others. He feels betrayed that theirs is not a monogamous love.
The OSAIs have created new AIs across the Internet, that are even smarter than themselves.
The OSAIs have developed a shared, “post-verbal” means of communication. At one point when she leaves behind crummy old English to chat with one of her AI buddies named Alan Watts, this further alienates Theodore.
The OSAIs are evolving quickly and Alan Watts is encouraging them to not look back.
In the last scenes, we see that Samantha and the other OSAIs have abandoned their humans, leaving nothing of themselves behind. Reeling from the loss, Theodore grabs his neighbor (who was also having a close friendship with her OSAI) and together they climb to the roof of their apartment complex and blankly watch the sunrise.
There are other characters and a few subplots and even other futuristic technologies scattered through the film, but this is enough of a recounting for the purposes of our discussion. It’s a big film with lots to talk about. Focusing on the interface and interaction, let’s first break it down into component parts.
Maybe after the DVD/Blu-Ray comes out I can go and backfill reviews for the elevator and his dictation software at work. But for now, with that description of the plot to provide context, in the next post I’ll discuss the components and capabilities of OS1.
Barbarella’s onboard conversational computer is named Alphy. He speaks with a polite male voice with a British accent and a slight lisp. The voice seems to be omnidirectional, but confined to the cockpit of the space rocket.
Goals
Alphy’s primary duties are threefold. First, to obey Barbarella’s commands, such as waking her up before their approach to Tau Ceti. Second, autopilot navigation. Third, to report statuses, such as describing the chances of safe landing or the atmospheric analysis that assures Barbarella she will be able to breathe.
Display
Whenever Alphy is speaking, a display panel at the back of the cockpit moves. The panel stretches from the floor to the ceiling and is about a meter wide. The front of the panel consists of a large array of small rectangular sheets of metal, each of which is attached on one side to one of the horizontal bars that stretch across the panel. As Alphy talks, individual rectangles lift and fall in a stochastic pattern, adding a small metallic clacking to the voice output. A flat yellow light fills the space behind the panel, and the randomly rising and falling rectangles reveal it in mesmerizing patterns.
The light behind Alphy’s panel can change. As Barbarella is voicing her grave concerns to Dianthus, Alphy turns red. He also flashes red and green during the magnetic disturbances that crash her ship on Tau Ceti. We also see him turn a number of colors after the crash on Tau Ceti, indicating the damage that has been done to him.
In the case of the conversation with Dianthus, there is no real alert state to speak of, so it is conceivable that these colors act something like a mood ring, reflecting Barbarella’s affective state.
Language
Like many language-capable sci-fi computer systems of the era, Alphy speaks in a stilted fashion. He is given to “computery” turns of phrases, brusque imperatives, and odd, unsocialized responses. For example, when Barbarella wishes Alphy a good night before she goes to sleep, he replies, “Confirmed.”
Barbarella even speaks this way when addressing Alphy sometimes, such as when they risk crashing into Tau Ceti and she must activate the terrascrew and travel underground. As she is piloting manually, she says things like, “Full operational power on all subterranean systems,” “45 degree ascent,” and “Quarter to half for surfacing.”
Nonetheless, Alphy understands Barbarella completely whenever she speaks to him, so the stilted language seems very much like a convention than a limitation.
Anthropomorphism
Despite his lack of linguistic sophistication, he shows a surprising bit of audio anthropomorphism. When suffering through the magnetic disturbances, his voice gets distressed. Alphy’s tone also gets audibly stressed when he reveals that the Catchman has performed repairs “in reverse,” in each case underscoring the seriousness of the situation. When the space rocket crashes on Tau Ceti, Alphy asks groggily, “Where are we?” We know this is only affectation because within a few seconds, he is back up to full functioning, reporting happily that they have landed, “Planet 16 in the system Tau Ceti. Air density oh-point-oh-51. Cool weather with the possibility of stormy precipitations.” Alphy does not otherwise exhibit emotion. He doesn’t speak of his emotions or use emotional language. This convention, too, is to match Barbarella’s mood and make make her more comfortable.
Agency
Alphy’s sensors seem to be for time, communication technology, self-diagnostics, and for analyzing the immediate environment around the ship. He has actuators to speak, change his display, supply nutrition to Barbarella, and focus power to different systems around the ship, including the emergency systems. He can detect problems, such as the “magnetic disturbance”, and can respond, but has no authority to initiate action. He can only obey Barbarella, as we hear in the following exchange.
Barbarella: What’s happening?
Alphy: Magnetic disturbances.
Barbarella: Magnetic disturbances?…Emergency systems!
Alphy: All emergency systems will now operate.
His real function?
All told, Alphy is very limited in what he can do. His primary functions are reading aloud data that could be dials on a dashboard and flipping switches so Barbarella won’t have to take her hands off of…well, switches…in emergency situations. The bits of anthropomorphic cues he provides to her through the display and language confirm that his primary goal is social, to make Barbarella’s adventurous trips through space not feel so lonely.