Abstract
Voice User Interfaces have become popular with the advent of Alexa, Google Home, Cortana and other commercial speech recognition interfaces; however, the privacy of the end users is compromised while using these interfaces in public. In addition, users can feel a bit awkward while using these interfaces with loud voice while they are outside their homes. Contextually, ‘Hypnosis’ and ‘Hypnotherapy’ have not been frequently applied as a way of human communications although the principle of suggestion induced behavior changes in an interesting approach to interact with machines. In this paper, GEORGIE a prototype AI was used to achieve a novel means of interaction inspired from the principles of hypnotherapy, which is a discrete interface ensuring that end-users’ privacy is not compromised. It is envisaged that people who prefer secret communication and interaction might love to use this hypnotic computer interface (hypCI). The hypCI would be the novel means of human robot interface (HRI) or human computer interface (HCI).
Keywords
Introduction
There are many examples of conversational inter-faces being used in use cases for hospitals [1], games for children [2, 3] and as assistive help for people with visual disabilities [4]. Gesture recognition has also been employed as means of effective human-robot interaction; the applications of this can be seen as an assistant to divers [5] and a means to interact with simulated environments [6]. There have also been several explorations on multimodal human robot interactions with speech being one of the modes of interactions coupled with gesture recognition and visual recognition. [7].
In the recent past, there has been a lot of concern over AI and Machine ethics has become an area, which has piqued a lot interest. Privacy is one of the focus areas in this domain, this is mainly due to the sporadic increase in the number of smart devices and most of them have voice as one of the modes of interaction.
Clinical psychologists have used hypnosis and Hypnotherapy for several decades for various ailments like anxiety disorders [8] and to treat addictions like smoking [9], alcoholism [10] and gambling [11]. Hypnotic triggers and cues are implanted by the hypnotist during sessions to achieve the desired results. A hypnotic suggestion is a very subtle form of interaction between the hypnotist and the subject, which is invisible to others like snapping of fingers or a whistle or other means [12–14]. Aim of this paper was to explore a new novel means of voice-based interactions with smart devices inspired from the principles in hypnotherapy and with the primary objective of making the interactions between humans and machines discrete.
Speech interfaces
Voice recognition software have been around for quite a while, the earliest ones go back to the 1960 s and 1970s when IBM introduced Shoebox [15] and Carnegie Mellon came up with Harpy [16]. In the 1990 s Dragon would introduce Dragon Dictate [17] the first voice recognition product aimed at consumers and post 2010 we would then see the onset of conversational interfaces like SIRI [18], Alexa [19], Google Home [20] and Cortana [21].
In the years gone by several advancements were made to voice recognition hardware and researchers have demonstrated various experiments to make the interaction with these interfaces seem intuitive and more human [22] and other researchers [23] through their work tried to make their speech interface understand context, other researchers have tried giving their interfaces personality [24]. These methods were helpful in making interaction with speech interfaces more efficient and thus made the robots sound more like humans. The privacy of speech interfaces is almost to none as it is evident that people can hear you talk to SIRI or Alexa in a public place [25].
Silent speech interfaces
A new form of discrete interaction with machines has been made by researchers keeping in mind the privacy issues of voice interfaces, these are known as silent speech interfaces. The origins of these interfaces are in sEMG and EEG based voice detection [26, 27], it was later discovered that inner voices and silent speech can be detected using sEMG electrodes using invasive [28] and non-invasive [29] techniques.
However, the main problem with sEMG voice recognition is not very efficient due factors like placement of electrodes and their design [30] hence it will take a few more years and development of better electronics to make a consumer targeted device.
Gesture recognition interfaces
Gesture recognition device can be classified into three types based on their input sensors namely proximity sensor based [31], sEMG based [32] and smart rings [33]. These devices partially solve the issue of privacy, even though the information is not shared with other people in the surrounding environment, however it gets evident that the users are interacting with their devices and the waving the gestures in air may seem a bit awkward for the users.
Human robot languages
Artificial languages have been made by humans for the purpose of fiction, the most popular ones are Klingon, Dothraki and others. The electronics behind speech recognition can only achieve a certain level of accuracy and speech recognition does not solve the problems of accent, keeping this point in mind and taking inspiration from fiction researchers are coming up with artificial languages to communicate with machines a good example of this is ROILA [34] which stands for Robot Interaction Language. While this method may help increase the privacy of the end users as people in the surroundings may not be aware of the language, however the interactions with the machine still do not remain discrete.
Hypnosis and hypnotic interfaces
There are a lot of misconceptions and stigma associated with hypnosis due to choreographed stage shows and movies; however, the topic has been heavily re-searched for several decades in areas of cognitive behavior, addiction research and neuroscience.
Process of hypnosis
A typical hypnotherapy process has two stages induction and suggestion. In the induction stage, the subject is asked to voluntarily perform a series of steps to put them in a trance like stage like focusing on one’s breathing or focusing on the sound of a metronome or visualizing guided imagery [35]. The result of this phase is to put the subject into a trance like state so that they can be made susceptible to accept behavioral changes. In the suggestion stage a series of triggers and suggestions are introduced into the subject psyche which would help achieve a targeted outcome like a change in behavior or quitting of certain habits like smoking or managing one’s anxiety. The post-hypnotic suggestions are useful to create the change in the subject.
Hypnotic triggers and suggestions have been a main area of focus, they are classified into different types depending on their nature namely verbal, non-verbal, intra-verbal and extra-verbal suggestions [36]. Non-verbal suggestions are very discrete like a set of gestures like snapping fingers, facial expressions or guttural sounds.
Hypnotic devices
Researchers have also made devices to induce post-hypnotic triggers, Spiegel [37] has come up with a smart watch that would trigger post-hypnotic suggestions. There have also been devices to induce hypnosis [38], induce a pre-hypnotic stage [39], mitigate treatments caused by therapeutic ICD shocks [40] and even anti-hypnotic devices for automobile drivers [41] as they seem to be focused on a single point for a long time which can induce hypnosis.
Research inferences
It was decided to classify the various solutions discussed in the above in terms of efficiency and privacy, the summary of the classification is showcased in the graph (Fig. 1) with level of privacy on the y-axis and efficiency of technology on the x-axis.

The below graph indicates how private is a mode of interaction versus how effective is the technology behind the mode of interaction.
Based on our preliminary research we have come up with the conclusions that the privacy of the user while using speech interfaces can be compromised and methods like silent speech recognition are still in its infancy and will take time to mature and be made as a commercial device.
Gesture recognition devices gives a partial privacy to the user but they are not a discrete form of interaction and the user may end up feeling awkward using gestures in a public place.
Hypnosis and hypnotherapy are an interesting form of interaction between the subject and the hypnotist and the suggestions are very discrete. There have been several hypnotic devices which have been used on humans, and being inspired from this research we are presenting a new form of discrete interactions between humans and machines named Hypnotic Computer Interfaces.
We have envisioned a new form of interaction between machines and humans known as Hypnotic Computer Interfaces abbreviated as hypCI. The key difference between hypCI and other interfaces is explained in the below diagram (Fig. 2).

The above diagram explains the key difference between a Hypnotic Computer Interface and any other interface.
We have defined hypCI as interfaces that utilize a discrete form of interaction between humans and machines borrowing the trigger-based interaction from clinical hypnosis. The triggers can be verbal, non-verbal like guttural sounds, gestures and facial expressions, intra verbal which are combination of two or more words and extra verbal triggers which are a combination of verbal and nonverbal triggers. In this form of interaction, the input can be from the humans and the feedback from the machine or vice versa.
In this paper, it was envisioning that an AI named GEORGIE (Guttoral Ekos Order Recognition and Gradually Indoctrinated Engine) which recognizes guttural sounds like a cough or a sniff meant as a non-verbal hypnotic trigger or suggestion. The user can be made to feel as if they are ‘hypnotizing’ the AI so that it can respond accordingly to the various triggers. For example: A user may program the AI to recognize a sniff as a yes or a cough as a no or book a cab home by making the noise of a vehicle.
System design
The trigger recognition engine’s working is explained in the system drawing (Fig. 3), depending on the type of trigger given by the user, the engine responds accordingly, if it’s an assigned trigger the engine proceeds with performing the specified task like telling the time or the current weather. If the user specifies it as a new trigger and assigns a specific task to it, the engine stores it and performs the specified task when triggered in the future, the user also can cancel a task by giving the voice trigger to stop the task, the engine ignores non assigned voices. A mobile application task flow of GEORGIE is also showcased in the below figure (Fig. 4).

The workflow diagram explaining the system design behind the AI GEORGIE.

The screen diagram explaining how the AI would work on a mobile interface. GEORGIE has a human like appearance to help the user connect more easily with it, the users can set the triggers and operate them accordingly.
As part of our study we are providing an extensive list of guttural commands (Table 1) and triggers which will enable the user execute specific set of tasks with GEORGIE. The users will also have the freedom to edit these commands on their will.
Extensive list of guttural commands with their respective tasks
Extensive list of guttural commands with their respective tasks
To evaluate the concept of Hypnotic Computer Interfaces, we prepared a prototype of the application on Axure RP. We used annyang javascript plugins to enable voice recognition in the axure prototypes. We prepared a pre-test questionnaire and a post-test questionnaire. A total of 30 users (M = 83% and F = 27%) took part in the evaluation of interface and the details of the questionnaires and the summary of findings are discussed in the below sections.
Pre-test questionnaire
Every user was made to fill up a pre-test questionnaire before performing the evaluations. They were asked questions regarding their preferred operating systems, the voice assistants they have used and their sense of privacy of using these speech interfaces in public places and their homes. The below figure (Figs. 5, 6 and 7) summarizes the user feedback from the pre-test questionnaire. Most of the users used Bixby as voice assistant, however, they are also almost equally frequent users of Google assistant, Cortana and Amazon Alexa. Most of the participants in this study were either academicians or students or developers. Maximum number of users are user of Ear-Pods. Most of the users are iPhone and IOS shabby.

Summary of Pre-test questionnaire used by the users (Please note that higher the size of the text means greater number of responses in this figure).

Perceived sense of privacy in a public place (1-Not at all private 7-Fully Private)

Perceived sense of privacy in homes (1-Not at all private 7-Fully Private)
Significant number of participants have felt that current voice-based interfaces doesn’t have privacy when using at public places [Privacy Public χ2 (2)=21.600; p < 0.001]. Therefore, voice-based interface users have privacy problem to use voice-based interface in public places. However, it almost undecided that whether voice-based interfaces are interfering with the privacy at home [PrivacyPublicχ2 (2)=7.400; p < 0.025].
The prototype for the evaluation was prepared on Axure with Annyang voice recognition plugins. The prototype had a mobile interface with instructions being provided beside for the user (Fig. 8.). The users were asked to set triggers for knowing the time, the weather and their heart rate to GEORGIE. They were then provided with a scenario in which they were asked to utter their set triggers and observe the feedback from the system. The prototype for the evaluation has only a set of 3 triggers (‘oh’ for heart rate; ‘roe’ for time and ‘aha’ for weather) but it can be programmed to include the whole list of voice commands presented in the above table (Table 1).

The prototype test interface deployed on a weblink for users to test along with instructions.
The users were asked to provide ratings of their experience while using the application on four different parameters i.e. Ease of use, Sense of Privacy, Comfort and Effectiveness of the mode of interactivity each on a scale of 1–7 with ‘1’ being strongly disagree and ‘7’ strongly agree. These parameters were adapted from [42–45]. The perceived accuracy, annoyance and habitability of voice control were also measured using modified SASSI (Subjective Assessment of Speech System Interfaces) scales [46]. Reliability of items were ensured by calculating the Cronbach’s alpha value. It was observed that alpha values of the accuracy, annoyance and habitability of voice control scales were 0.87, 0.88 and 0.71 respectively. All of these Cronbach’s alpha value is more than the value of 0.70 which is the minimum requirement to say a scale as reliable. Hence, all these modified SASSI scales were reliable.
Results
The average ratings for each of the measure factors are showcased in the below figures along with the standard errors (Please see the Fig. 9). We also performed a chi square test to determine if there is a major difference between expected response and anticipated results. Since most of the users gave rating s above the neutral value (4.00) in the seven-point Likert scale, we decided to calculate the frequency of ratings above or equal to the mean to determine if the mode of interaction was really effective.

The above graph showcases average ratings along with the standard error for Ease of Use, Privacy, Awkwardness and Effectiveness.
The summary of these findings showcased in Fig. 10 indicates that the most of the users felt the solution was comfortable, easy to use (EoU) and effective [Comfortχ2 (2)=19.400; EoUχ2 (2)=15.207; EFχ2 (2)=38.600; p < 0.001]. Other studies also support that if the interface is easy / comfortable to use and effective then there is a chance of system acceptance [44, 45]. Significant number of users also felt the mode of interaction was private [Privacyχ2 (2)=30.200; p < 0.001] and did not feel awkward while using them. Few users pointed out in the interview that “the idea of hypnotizing the AI made him feel empowered” and “he felt in control of the AI and less fearful”. Thus, the proposed interface solved the problem of low privacy of current voice-based user interface. Therefore, user might accept this kind of voice-based hypCI interface in future.

The above graph showcases the percentage of ratings equal to or above the mean value.
It was observed that mean values of perceived accuracy and habitability of voice control is greater than 03 (means neutral) on 5-point Likert scale. On chisquare test [Accuracyχ2 (2)=7.800; Habitabilityχ2 (2)=18.600; p < 0.05] it was observed that significantly most of the participants responded about the system as more accurate (57%) and habitual (67%) (Please see Fig. 11 and Fig. 12). On the other side, average score for annoyance was less than 03 (means neutral) on 5-point Likert scale and a significant number of participants (67%) perceied the interface as less annoying [Annoyanceχ2 (1)=10.800; p < 0.05] (Please see Figs. 11 and 12). All these results signify that users were preceived the proposed system as more accurate and users were habituated with the voice user interface, however, the interface was not annoying.

The above graph showcases average ratings along with the standard error for perceived accuracy, annoyance and habitability of voice control.

The above graph showcases the percentage of ratings above the value 3 (Neutral) in 5-point Likert Scale.
The initial user feedback for the experimental prototype was positive. Hence, proposed interface could be adopted and applied for various purposes such as secret communication in defense, colour blind people who hesitate to vocally ask for assistance from voice assistance like Alexa or Google assistant.
It is possible to work on improving the prototype by adding more functionalities and make it less erroneous. In the future, we would like to research more on extra verbal and intra verbal suggestions to interact with smart devices, it would also possible to explore how mitigating fear of AI using this method based on the feedback of users. It is anticipated that the proposed mode of interaction would be highly effective for Smart Glasses as they have voice as a primary mode of interaction, hence, studies would be conducted with a smart glass prototype in near future.
Footnotes
Acknowledgments
We would love to thanks all participants (users) who provided their valuable feedbacks during this study without which this paper would not possible.
