Abstract
There has been a significant amount of work implementing systems for algorithmic composition with the intention of targeting specific emotional responses in the listener, but a full review of this work is not currently available. This gap creates a shared obstacle to those entering the field. Our aim is thus to give an overview of progress in the area of these affectively driven systems for algorithmic composition. Performative and transformative systems are included and differentiated where appropriate, highlighting the challenges these systems now face if they are to be adapted to, or have already incorporated, some form of affective control. Possible real-time applications for such systems, utilizing affectively driven algorithmic composition and biophysical sensing to monitor and induce affective states in the listener are suggested.
Affective algorithmic composition (AAC) is a field which combines computer aided composition and emotional assessment. The field presents an interdisciplinary challenge to music psychologists, composers, and computer music researchers. This article presents an overview and summary of AAC systems and their uses. In order for a system to be considered AAC it should include some implementation of affect, usually an affective target. The affective target might be a single affective descriptor, a combination of descriptors, or a position on a dimensional model. This field has been termed affective algorithmic composition in recent work (Kirke & Miranda, 2011; Kirke, Miranda, & Nasuto, 2012; Mattek, 2011; Williams, Kirke, Miranda, Roesch, & Nasuto, 2013), though this term could equally encapsulate earlier work such as affectively “driven” or “flavoured” algorithmic composition systems (Legaspi, Hashimoto, Moriyama, Kurihara, & Numao, 2007; Wallis, Ingalls, & Campana, 2008; Wallis, Ingalls, Campana, & Goodman, 2011), or systems using artificial intelligence techniques to inform the affective trajectory (Kim & André, 2004).
AAC systems are distinct from traditional algorithmic composition systems that do not consider an intended affective trajectory in the generated material: in AAC, the algorithm is always informed by an intended affective response. This affective targeting can be enhanced by the possibility of real-time adjustment of the algorithm based on affective metering in a feedback loop. Thus, the development of AAC systems can be more complex in terms of implementation, and evaluation, than other types of algorithmic composition.
The potential applications for a successful AAC system are many. A composer might use AAC to deliberately attempt to parametrically control the listeners’ mood. Or, a listener might use an AAC system combined with an affective metering system (biophysical, self-reported, or otherwise) to create responsive music that takes into account their own emotional state (e.g., for meditation purposes). Thus, a musical “layman” without the ability to compose or perform might use AAC to create novel and emotionally satisfying music.
First steps towards medical applications for such work have been made in research measuring the effect of background music on the neurophysical response of children with attention deficit disorder (Rebollo, Hans-Henning, & Skidmore, 1995; Strehl et al., 2006) and most recently in a brain-computer control system allowing users to control musical parameters via electroencephalography (EEG, often referred to as the “brain cap”) (Eaton & Miranda, 2013; E. Miranda & Brouse, 2005; E. R. Miranda, Sharman, Kilborn, & Duncan, 2003). In one such example, E. R. Miranda, Magee, Wilson, Eaton, and Palaniappan (2011) describe the evaluation of a pilot brain-computer musical interface allowing a patient with Locked-in syndrome to control amplitude and other musical parameters via EEG for the purposes of music therapy and palliative care at the Royal Hospital for Neuro-disability in London, UK. This type of system could be further expanded to affective control by utilizing the growing body of work correlating EEG responses to affect (Chanel, Kronegg, Grandjean, & Pun, 2006; Lin et al., 2010; Schmidt & Trainor, 2001) as part of the control mechanism for an AAC system.
Emotional assessment in empirical work often makes use of recorded music (Gabrielsson & Juslin, 2003; Gabrielsson & Lindström, 2001; Wedin, 1969, 1972) or synthesized test tones (Juslin, 1997; Scherer, 1972) to populate stimulus sets for subsequent emotional evaluation. There are some reported difficulties with such evaluations with specific regards to measurement of emotional responses to music. For example, some research has found that the same piece of music can elicit different responses at different times in the same listener (Juslin & Sloboda, 2010). If we consider a listener who is already in a sad or depressed state, it is quite possible that listening to “sad” music may in fact increase listener valence. One challenge then for truly affective algorithmic composition systems is to be able to respond to the initial, and potentially dynamically changing, affective state of the listener in real-time.
Another challenge for such evaluations is that music may intentionally be written, conducted, or performed in such a manner as to be intentionally ambiguous. Indeed, perceptual ambiguity might be considered beneficial as listeners can be left to craft their own discrete responses (Cross, 2005). The breadth of analysis given to song lyrics gives many such examples of the pleasure listeners can take in deriving their own meaning from seemingly ambiguous music. There is an opportunity to adapt biophysical sensing of emotions to the control of AAC systems, such that they can be used to respond adaptively to individual listeners’ affective states, or indeed to “affective interpretations” of otherwise ambiguous musical structures.
This article presents an overview of existing work, including systems that combine composition with affective performance (Gabrielsson, 2003; Gabrielsson & Juslin, 1996; Palmer, 1997), with a view to further development. We first define the terminology that will be used, including the various psychological approaches to documenting musical affect, and the musical and acoustic features that such systems utilize. Case studies of systems that address AAC using different emotional models and different generative algorithms are presented. Systems covering the largest number of features are then outlined, and compared by underlying affective models, emotional correlates, and use of musical feature-sets.
Background: Terminology and concepts
This section introduces the terminology and the concepts that form the basis for this assessment of AAC systems. We do not aim to be exhaustive, but rather to provide the reader with an overview of the psychological phenomena that often inform the development of such systems. Interested readers can find more exhaustive reviews on the link between music and emotion in Scherer (2004) and the recent special issue in Musicae Scientiae (Lamont & Eerola, 2011).
Literature concerning the psychological approaches to musical stimuli broadly documents three types of emotional response, each increasing in duration:
Emotion: A short-lived episode; usually evoked by an identifiable stimulus event that can further influence or direct perception and action. It is important to note here although listeners may experience emotions in isolation, the affective response to music is more likely to reflect collections and blends of emotions.
Affect/subjective feeling: Longer than an emotion, affect is the experience of emotions or feelings evoked by music in the listener.
Mood: Longer-lived again, mood is a more diffuse affective construct; usually latent and indiscriminate as to eliciting events; mood may influence or direct cognitive tasks.
Particularly relevant to AAC is the distinction between emotion and subjective feeling (the part of emotion which is consciously accessible to the person experiencing the emotion), whereas the other types of affective response are not necessarily available to conscious report and have practical consequences over components of emotion (such as motor expression or action tendencies). In any case, measuring these responses is difficult – the former components may or may not be consciously perceived, or correctly reported by the person experiencing the emotion. Bodily symptoms alone are not sufficient to evoke and consequently allow for the reporting of emotions (Schachter & Singer, 1962). A full treatment of the underlying mechanisms at play in accounting for evaluation of musical emotion is given by Juslin and Västfjäll (2008).
Emotional models
There are two main types of emotional models used in relation to affective analysis of music – categorical and dimensional models. Categorical models use discrete labels to describe affective responses. Dimensional approaches attempt to model an affective phenomenon as a set of coordinates in a low-dimensional space (Eerola & Vuoskoski, 2010). Discrete labels from categorical approaches (for example, mood tags in music databases) can often be mapped onto dimensional models, giving a degree of convergence between the two. Neither are music-specific emotional models, but both have been applied to music in many studies (Juslin & Sloboda, 2010). More recently, music-specific approaches have been developed (Zentner, Grandjean, & Scherer, 2008). Implementations of dimensional approaches in AAC systems are very popular (though various systems use categorical approaches – see the section on case studies in this article for examples of each). The circumplex dimensional model, for instance, describes the semantic space of emotion within two orthogonal dimensions, valence and arousal, in four quadrants (e.g. positive valence, high arousal). This space has been proposed to represent the blend of interacting neurophysiological systems dedicated to the processing of valence (pleasure–displeasure) and arousal (quiet–activated) (Russell, 2003; Russell & Barrett, 1999). The Geneva Emotion Music Scale (GEMS) describes nine dimensions that represent the semantic space of musically evoked emotions (Zentner et al., 2008), but unlike the circumplex model, no assumption is made as to the neural circuitry underlying these semantic dimensions (Scherer, 2004). GEMS is a measurement tool to guide researchers who wish to probe the emotion felt by the listener as they are experiencing it, and might therefore provide a useful model for future AAC development with real-time and adaptive applications.
Perceived vs. induced
Zentner, Grandjean, and Scherer (2008) carried out experiments that examined the differences in felt and perceived emotions. Their conclusion highlights the difficulty faced by AAC systems: “Generally speaking, emotions were less frequently felt in response to music than they were perceived as expressive properties of the music” (Zentner et al., 2008, p. 502). This distinction is important when considering AAC systems, and has been well documented, see for example Västfjäll (2001), Gabrielsson (2001a), and Vuoskoski & Eerola (2011), though the precise terminology used to differentiate the two varies widely, as summarized in Table 1. Perhaps unsurprisingly, results tying musical parameters to induced or experienced emotions do not often provide a clear description of the mechanisms at play (Juslin & Laukka, 2004; Scherer, 2004), and the terminology used is inconsistent.
Synonymous descriptors of ‘Perceived/Induced’ emotions that can be found in the literature. For detailed discussion the reader is referred to (Gabrielsson, 2001a; Kallinen & Ravaja, 2006; Scherer, 2004).
Research investigating emotional induction by music
In the context of AAC, an induced emotion would be an affective state experienced by the listener, rather than an affect which the listener understands from the composition — by way of example, this would be the difference between listeners reporting that they have “heard sad music” rather than actually “felt sad” as a result of listening to the same.
For a more complete investigation of the differences in methodological and epistemological approaches to perceived and induced emotional responses to music, the reader is referred to Gabrielsson (2001a), Scherer, Zentner, and Schacht (2002), and Zentner, Meylan, and Scherer (2000).
Types of algorithmic composition
Musical feature-sets are often used as the input for algorithmic composition systems. Algorithmic composition (either computer assisted or otherwise) is gradually becoming a well-understood and documented field (N. Collins, 2009; E. R. Miranda, 2001; Nierhaus, 2009; Papadopoulos & Wiggins, 1999), though new techniques are being constantly developed, from incremental refinements in parameterization and selection algorithms through to novel artificial intelligence (Moroni, Manzolli, Von Zuben, & Gudwin, 2000) or genetic algorithm techniques (Gartland-Jones, 2003; Horner & Goldberg, 1991; Jacob, 1995).
Rowe (1992) describes three methodological approaches to algorithmic composition: generative, sequenced, or transformative. Generative systems use rule-sets to create musical structures from control data. Selective filtering of notes generated by a random or semi-randomized function, as in Hiller and Isaacson’s Iliac Suite for String Quartet (1957), would be considered generative. The generative processes employed can also include a significant amount of variation, including, for example, composer defined, non-linear, or randomized functions (Harley, 1995). Sequenced systems take pre-defined “chunks” of music and order them according to some rule-based input selection. This can be adapted to affectivity by using chunks which have been pre-rated in terms of emotional response (in other words, using the desired emotional response as the input selection). Performance elements of a sequenced system may be further varied (for example, tempo or rhythmic variations introduced) to target affective response. Mozart’s Musikalisches Würfelspiel (Nierhaus, 2009) can be considered an example of a sequenced system, re-ordering pre-composed sections of music based on the input signal given by the dice roll. In such systems, the composer has a clearer influence on the musical output than in a purely generative system, though this might be considered a hindrance by composers seeking entirely novel compositional output, or, indeed by non-musically literate listeners using AAC for applications such as in the meditative examples given earlier. Transformative systems use existing material as the source. Here, one or more transformations are applied to the input material in order to yield related material at the output stage. A simple inversion might be considered an example of a transformative process. More complicated transformative processes allow some systems to “ape” existing styles by process of deconstruction, analysis, and recombination, as in the Experiments in Musical Intelligence work of Cope (1989, 1992).
Structure in AAC: Defining “musical features”
Musicologists have a long-established, though often evolving, grammar and vocabulary for the description of music, in order to allow detailed musical analysis to be undertaken. However, for the purposes of outlining and evaluating AAC systems, an in-depth knowledge of analysis and musical structure is not necessary. Therefore we will only briefly introduce the musical and/or acoustic features used in some AAC systems here.
Melodic, harmonic, and rhythmic content alone will not by default create recognizably musical structures (Bent, 1987). Musical themes emerge as temporal products of these features (Hanninen, 2012) – melodic and rhythmic patterns, phrasing, harmony and so on.
Figure 1 provides an overview of the inputs and outputs an algorithmic composition system might use in order to produce an affective output.

Overview of basic affective algorithmic composition system, including optional performance system. A minimum of three inputs are required: algorithmic compositional rules (generative, or transformative), a musical (or in some cases acoustic) dataset, and an emotional target. Emotional correlates, determined by literature review of perceptual experiments, can be used to inform the generative or transformative rules in order to target specific affective responses in the output.
Musical features as emotional correlates
Table 2 presents a non-exhaustive overview of existing research that correlates emotions with musical feature(s). Structural rules, which might be implemented in an AAC system, are the main focus of this review. It is acknowledged that an emotion might be a response to several independent musical features (Hevner, 1935; Levi, 1978), and that notwithstanding, the studies presented often use different stimuli (e.g., speech, chord progressions, isolated tones, original compositions, existing compositions, etc.). For these reasons, the systems in Table 2 are presented chronologically. The musical features and emotional correlates implemented in 16 systems available at the time of writing are summarized. These systems can be further categorized according to the emotional model used. Much of the work uses the 2-Dimensional or circumplex model of affect (valence and arousal). Other common approaches include single bipolar dimensions (such as happy/sad), or multiple bipolar dimensions (happy/sad, boredom/surprise, pleasant/frightening etc.). Finally, some models utilize categorical, free-choice responses to musical stimuli to determine emotional correlations. Neither the single or multiple bipolar dimensional approaches, nor the free-choice responses, are exclusive of the 2-Dimensional model. Many of these approaches share commonality in the specific emotional descriptors used:
11 systems use the 2-Dimensional model of affect as the basis for evaluation of emotional correlations
4 systems use multiple dimensions as the basis for evaluation of emotional correlations
3 systems use single bipolar dimensions (fear/anger, happy/sad, joyful/melancholic)
3 systems use single emotional correlations (tension, arousal, surprise)
3 systems use “free choice” profiling of emotional responses
A summary of musical feature-set and emotional correlates implemented in existing systems.
Existing systems and feature-sets
This section introduces a feature based overview of existing systems, outlining the data sources used. The range of musical feature-sets incorporated by existing systems is then analysed and discussed. Three specific case studies are then presented. Finally, the AAC systems that include the largest number of features are presented in further detail, including the affective approach (emotional model and correlates), and the musical feature-sets used.
System features
Existing systems using algorithmic composition to target affective responses can be categorized according to their data sources (either musical features, emotional models, or both), and by their approach to the composition process. The process can be categorized in a broadly bipolar fashion as follows:
Compositional/Performative. Does the system include both compositional processes and affective performance structures? Compositional systems refer synonymously to structural, score, or compositional rules. Performative rules are also synonymously referred to by some research as interpretive rules for music performance. Certain styles of music might make compositional use of features that would otherwise be considered performative features (for example, micro tuning, or expressive timing). The distinction between structural and interpretive rules might be considered as differences that are marked on a score (for example, dynamics might rely on a musician’s interpretive performance, or be structures sequenced with expressive articulations that are still a part of the compositional intent). For a fuller examination of these distinctions, the reader is referred to Gabrielsson & Juslin (1996) and Gabrielsson & Lindström (2001).
Perceived/Induced. Does the system target affective communication, or does it target the induction of an affective state?
Adaptive/Non-adaptive. Can the system adapt its output according to its input data (whether this is emotional, musical, or both)?
Generative/Transformative. Does the system create output by purely generative means, or does it carry out some transformative/repurposing processing of existing material?
Real-time/Offline. Does the system function in real-time? Various AAC end-uses might necessitate real-time processing.
A summary of the use of these approaches amongst existing systems is given in Table 3.
A summary of features (where known or implied by literature) employed by existing systems for affective algorithmic composition.
None of the systems listed in Table 3 specifically target affective induction through generative or transformative algorithmic composition in real-time. Issues of induction, ambiguity, and biophysical measurement all add to the complexity of such a system. This gap presents a significant area for further work in order to create AAC systems that can respond to a listener’s existing affective state (for example in the “meditation” or video game music use-cases).
Musical features and existing systems
The systems outlined in Table 3 utilize a variety of musical features. Perceptually similar and synonymous terms abound in the literature and thus deriving a ubiquitous feature-set is not a straightforward task. Though the actual descriptors used vary, a summary of the major musical features found in these systems is provided in Table 4. These “major” features are derived from the full corpus of terms by a simple Verbal Protocol Analysis (VPA). VPA is a useful technique for analysing the instances of descriptors by grouping terms according to similarity of meaning and then calculating a measure of the significance of each group. The most prominent features are used as headings in Table 4, with an implied perceptual hierarchy.
Number of generative systems implementing each of the major musical features as part of their system. Terms taken as synonymous for each feature are expanded in italics. Major terms are presented left to right in decreasing order of number of instances. Minor terms are presented top to bottom in decreasing order of number of instances, or alphabetically by first word if equal in number of instances.
The largest variety of sub-terms comes under the “Melody (pitch)” and “Rhythm” headings, which indicate the highest level of perceptual significance in terms of a hierarchical approach to musical feature implementation. With the addition of timbre, the upper level of the emerging ontology is revealed, as illustrated in Figure 2. Below pitch we find harmony, melody, and so on. Contributing to rhythm we find one of the most unequivocal descriptors, tempo, which seems to have no synonymous features in the corpus. Whilst “mode” and its synonyms are nominally the most common in the pitch category, the results also show a lower number of instances of the word “mode” or “modality” than “pitch” or “rhythm,” suggesting those major terms to be better understood, or perhaps, more appropriate to the systems surveyed. Whilst “timbre” appears only three times in the group labelled “Timbre,” which includes five instances of noise/noisiness and four instances of harmonicity/inharmonicity, timbre has been used as the heading for this umbrella set of musical features given the particular nature of the other terms included within it (timbre provides commonality between each of the terms in this heading). Harmonicity is a component which suggests an inter-relationship between the higher levels, pitch content could certainly influence the harmonicity, and noisiness of a musical timbre. However, the majority of literature identifies harmonicity as a contributory timbral component, rather than a primarily pitch-derived one. A similar assumption might be made about dynamics and loudness, where loudness is in fact the most used term from the group, but the over-riding meaning behind most of the terms can be more comfortably grouped under dynamics as a musical feature, rather than loudness as an acoustic (or psychoacoustic) feature.

Some musical features shown in a diagram of the emerging ontology of features used in affective algorithmic composition systems. A secondary inter-relationship between harmonicity and pitch is designated by the dashed line (harmonicity is considered to be a contributing component of timbre, but has some relationship to pitch content).
Under the “Melody (pitch)” label, there could be an eighth major division, pitch direction (with a total of 8 instances in the literature, comprising synonymous terms such as melodic direction, melodic change, phrase arch and melodic progression), implying a feature based on the direction and rate of change in the pitch.
Case studies
Case studies examining the details of three specific systems are presented here to illustrate the differences in approach, both generatively and affectively, undertaken by different types of AAC. These systems have been selected in order to illustrate the wide variety in algorithmic and affective approaches to AAC, with an example of a sequencing algorithm, a transformative algorithm, and a generative algorithm, each using different emotional approaches and musical feature-sets. Each system is briefly introduced before an overview of the program structure is given. The implementation of musical features and approach to emotion is then given, followed by a comparative discussion of the generative and affective potential of each case.
Case study 1: Moodtrack
Moodtrack (Chung & Vercoe, 2006; Vercoe, 2006) is a system for arranging music based on an intended emotional target. Moodtrack is designed with affective sound-tracking for film in mind, but could potentially be adapted to other AAC applications. Moodtrack uses a high level MIR scheme to extract an emotional trajectory for metadata matching to material in existing soundtracks.
The source material is manually segmented into phrases and saved as a library. The library is then evaluated affectively in a two stage process by musical experts and online gamers. Emotional annotations are then assigned to each segment in the library. An emotional trajectory consisting of desired emotional contour and timing data is supplied as the input, and the MIR scheme then matches appropriate sections from the library, compiling segments as a new score.
The emotional approach taken by Moodtrack is essentially categorical, based on the adjective cycle devised by Hevner (1937) with the addition of a 100-point user-reported value for intensity of emotion, and some additional descriptors. The system utilizes six bi-polar features from Hevner: mode, tempo, rhythm, melody, harmony, and register, with the addition of four further features: dynamics, timbre, density, and texture.
The generative potential of Moodtrack is somewhat limited as precomposed segments are assembled to match the emotional cues and time co-ordinates supplied by the user. However, the generative potential could be increased by using larger databases of source material, and reducing the segment size. The system is emotion-driven, rather than directly evaluating the affective characteristic of the output. The affective potential is significant, with a large number of emotional categories and an intensity vector to allow listeners to evaluate segments in a comprehensive manner, though even with the extension of emotional categories by means of additional descriptors, the use of a predetermined lexicon for emotional responses might be considered limiting.
Case study 2: EDME
The Emotionally Driven Interactive Music System (EDME) (Oliveira & Cardoso, 2007; Ventura, Oliveira, & Cardoso, 2009), uses an artificial intelligence approach to express an emotional state via a parameterized set of musical features. The authors also describe how the system might be used as a tool for composition, similarly to Moodtrack, for purposes of soundtrack generation for computer games or interactive media.
The system uses an emotional descriptor as its input source. The algorithm is based on a transformation of pre-composed and automatically evaluated scores in MIDI format. The source material is segmented and musical features in each segment are extracted, evaluated, and assigned a weighting. This part of the system is similar to the approach taken by Moodtrack to generate a database of source material, but the material in the database is classified in emotional correlation to a 2-D arousal and valence model by means of a regression, and the segments are subsequently passed through a transformation stage having been selected according to the emotional target entered in 2-D (arousal and valence). EDME also varies procedurally in that it uses a multi-agent system at the generation stage to compile the score.
Five musical features are used by EDME: pitch, rhythm, silence, loudness, and instrumentation. These features are weighted in each segment according to their prominence. This weighting is employed in the automatic segmentation of source material, which is segmented according to the weighting given to note onsets (prominent onsets define new segments). This is unlike the approach taken by Moodtrack which utilizes ‘hand’ segmentation of source material.
The generative potential of EDME is, theoretically at least, somewhat larger than that of Moodtrack, by virtue of the addition of multi-agent transformations to the algorithm. However, EDME material has yet to be evaluated in listening tests, and the challenge of creating affectively congruent transformation algorithms is not addressed in the documentation. Ultimately, the generative potential of any AAC system based solely on transformation algorithms is limited. The authors of EMDE give interesting possible applications in the form of real-time adaptation of their system to biophysical control, and point to possible clinical use, but have yet to evaluate these applications in any published trials.
Case study 3: Roboser/EmotoBot
Roboser (Manzolli & Verschure, 2005) uses a machine learning system, EmotoBot, to control the generation of streams of MIDI data in a real-time composition. The control data is collected by robot sensors and modulated through a neural network before being mapped onto various musical features by the composition engine. The robot moves using light and collision sensors to help it to explore an area and try to avoid obstacles. The machine learning system stores positive outcomes in memory and tries to repeat the behaviour that led to them. This data is used to give an indication of affective state for the robot. The data is then used as a control to select or transform a score from a small pool of seed data in real-time as the robot goes through its behavioural patterns.
The approach to emotion in Roboser, a combination of machine learning and neural mapping, is significantly different to that taken by Moodtrack (which used a modified Hevner adjective cycle) or EDME (which used the 2-D circumplex model of valence and arousal). Various affective states were mapped to MIDI note number and rhythmic pattern transformations in the composition engine. The features include timbre (a range of four MIDI voices), velocity, tempo, and pitch. The mapping uses a variety of preset values for these features, for example, when the robot is in an exploratory state, a selection of one in four rhythmic patterns are possible. In the case of a collision, the MIDI voice would change, and tempo would increase. Movement and affective state is thus represented by the score created by the composition engine.
The results presented by the creators of Roboser suggest a wide variety of generative potential from an initially small number of seed constraints, particularly with the transformation of material held in long term memory by the system. Roboser was shown to generate new musical structures in response to different affective states, but the overall affective potential of the system is yet to be evaluated at present. Roboser provides a novel approach to AAC which might be applicable to use-cases where the goal is the creation of affectively charged music without the need for any particular musical expertise in the end user. Such applications include communication systems for patients with severe disabilities, and the generation of therapeutic music. There is also potential to expand this system (and others like it) with a larger neural network as an affective control signal.
Feature sets in other systems
Of the 32 systems outlined in Table 3, seven systems cover at least six of the eight specified feature-sets. The Computational Music Emotion Rule System (CMERS) (Livingstone, Mühlberger, Brown, & Loch, 2007), combines a composition and performance system developed through analysis-by-synthesis. In CMERS, a rule system pairing intended emotional response and musical features was used to create MIDI filters. MIDI data from an existing score is passed through these filters to shape a synthesized performance that is then analysed by listener testing to confirm the intended and perceived emotional responses to the output. CMERS includes micro-feature deviations to generate “humanized” performances, further modifying the performance from the score. Musical features incorporated include mode, pitch height, tempo, loudness, articulation, and micro-timing deviations, and emotional correlations have been found with angry, bright, contented, despair, etc.
Some systems are already moving towards direct biophysical control. Kirke and Miranda (2011a) describe a system which used frontal asymmetry measured from the electroencephalography of listeners to drive an AAC system with expressive performance rendering. Again, this system utilizes the arousal-valence dimensional model of emotion. The musical features generated include a hierarchical structure of random motifs (a generative approach) with an algorithmically generated left-hand piano accompaniment (a transformative approach). Hence this system could be considered simultaneously transformative, generative, compositional, and performative.
The Emotional Music Synthesis (EMS) system, (Wallis et al., 2008, 2011), utilizes a generative system controlled by an intended valence and arousal rating. This system gives real-time, parameterized emotional control over musical feature-sets which are correlated with the 2-D emotion space. Listener responses suggest that timing and volume features are correlated to arousal, whilst tonality and timbre features are correlated to valence, with mode being the most perceptually important feature for valence, and density the most important for arousal.
SiMS, an adaptive music engine which utilizes reinforcement learning and agent interaction (Le Groux & Verschure, 2010) similarly parameterizes musical features as control signals, though for a single emotional correlate, tension. This system introduces a perceptual synthesis engine, manipulating acoustic features (amplitude envelope, inharmonicity, noisiness, and harmonic ratios), as well as utilizing monophonic transformative processing of rhythm, pitch, register, and dynamics, and polyphonic chord generation.
These summaries give some impression of the scale and utility of various approaches taken by AAC systems to date.
Conclusions
Affective Algorithmic Composition (AAC) is a proposed umbrella term that includes any system for composition designed to respond to an affective target and/or to create an affective response in the listener. Such systems have various possible uses, including responsive sound-tracking for film or video games, affective communication aids, and therapeutic music generation. For example, in the world of feature film, sound-tracking is often used to enhance the emotional impact of a scene (and could conceivably change the emotional context of a scene completely). Music created for the video game industry includes the additional complication that the accompanying score might be required to adjust on-the-fly in response to unpredictable narrative changes. Generative AAC systems for the creation of novel material can be extremely useful in this context. Progress in neurophysical monitoring (including brain-computer interfacing such as EEG ‘brain cap’ readings) suggests that in the future, a robust system combining affective metering via neurofeedback and real-time music generation might be used to help people who struggle to communicate verbally – patients who suffer from Asperger syndrome, autism, or Locked-in syndrome, for example. In a fully developed therapeutic system, these biophysical measures could be used to monitor the listeners’ affective states and, in response, create appropriate musical structures in near real-time. As well as a communication aid for patients who might otherwise struggle, this type of system could be used to help move listeners through different affective states (for example to help with depression, or to aid relaxation). Thus, there is an opportunity to develop neurofeedback-derived control over musical features in response to individual affective responses in such a system. One approach would be the development of an affective matrix that comprises parameterized feature trajectories based on existing state and target state, controlled in real-time by neurofeedback. In such a system, a generative approach would offer clear advantages over affective remixers, interpolators and the like by sustaining novelty in the output and countering listener fatigue in longer sessions.
A generic overview of AAC systems has been presented, including a basic vernacular for classification of such systems by feature-set, seed material, and control data. The control data is often an affective target, trajectory, or contour, and most systems make use of some restrictions or modifications to the seed material as a starting point for the resulting musical structures that are generated. The musical feature-sets and emotional approaches employed by these systems vary quite significantly, as found in case studies of existing AAC systems.
Algorithmic composition approaches are normally sequenced, transformative, or generative in nature. A case study of an AAC system using each type of algorithm has been presented here. The first case study, Moodtrack, used a categorical model of emotion with a sequencing approach to composition to re-assemble existing segments of music based on documented emotional responses. The generative potential of such a system is somewhat limited by the size of the seed database, though novelty in soundtrack generation is less of an issue with systems for creation of affective film sound-tracking (the use for which Moodtrack was originally designed). This type of system would probably be less appropriate for the video game world, where it is possible that a given affective target might yield very similar material from the seed database for a long time, depending on the state of gameplay. A second case study found that EDME employed a dimensional approach to affect, targeting arousal and valence with a transformative algorithm. A third case study, Roboser/EmotoBot showed how a generative approach could use a small number of seed patterns to generate a large amount of musical structure from a small seed pool in response to a neural network in real-time. A significant amount of novelty was documented by the creators of this system despite a relatively small amount of seed material. Of all the case studies, this example is perhaps best suited to the meditative or therapeutic applications mentioned earlier. At the time of writing there was a lack of specific affective evaluation by listener testing of the system, though both the affective and generative potential of this system is significant.
Which musical features are most commonly implemented in AAC systems?
Modality, rhythm, and pitch are the most common features found in AAC systems, with 30, 29, and 28 instances, respectively, found in the literature. These musical features indicate an implicit hierarchy with, for example, pitch contour and melodic contour features making a significant contribution to the instances of pitch features as a whole.
The approaches to musical features taken by AAC systems varies widely, as they do in “normal” algorithmic composition systems, with seemingly little agreement as to which features are essential, desirable, dispensable, and so on.
Which emotional models are employed by AAC systems?
Other dimensional approaches exist, including implementations of categorical models, but the 2-D (or circumplex) model of affect is by far the most common of the emotional models implemented by AAC systems. Multiple and single bipolar dimensional models are employed by the majority of remaining systems. The existing range of emotional correlates, and in some cases the bipolar adjective scales used, is not necessarily evenly spaced in the 2-D model. GEMS specifically approaches musical emotions, allowing for a multidimensional approach (Fontaine, Scherer, Roesch, & Ellsworth, 2007) and providing a categorical model of musical emotion with nine first-order and three second-order factors. GEMS would provide an opportunity for emotional scaling of parameterized musical features in an AAC system but none of the surveyed AAC systems have utilized this approach to musical emotion. Indeed, affective evaluation in the surveyed AAC systems is sparse. There is a significant amount of further work in such evaluations. A system for the real-time, adaptive induction of affective responses by algorithmic composition (either generative or transformative), including the affective evaluation of music by measurement of listener responses to such a system also remains a significant area for further work. If a multidimensional approach based on GEMS was taken by an AAC system, a useful starting point would be the selection of musical features with emotional correlates that are as dissimilar as possible (that is, as spatially different in the emotion-space) to aid an initial affective evaluation.
This review of the link between musical features and emotional responses in AAC is being used to inform work on a new research project, Brain Computer Musical Interface for Monitoring and Inducing Affective States (BCMI-MIdAS). One of the project goals is to provide a system for parameterized control of emotional responses within an automatic composition model, to induce specific affective states automatically and adaptively. Both generative and transformative engines could be used as the basis for this type of system, though transformative algorithms would lend themselves particularly well to the measurement of the effects of relative changes in musical structure (such as changes in modality or melodic features that have no easy “baseline” reference measure). A generative approach would lend a greater degree of possible novelty to such a system, and perhaps also increase its usefulness in meditative or therapeutic applications, and video-game applications, where duration is not readily specified in advance. A combined generative and transformative approach would be a useful starting point for anyone interested in developing the “next generation” of AAC system with biophysical control for induction of emotion by music in real-time.
Footnotes
Funding
This work was supported by the Engineering and Physical Sciences Research Council, [grant numbers EP/J003077/1, EP/J002135/1].
