Thursday, July 23, 2015

10 tips for 10-minute presentations

One of the things I love most in life is helping people improve their academic talks. My advice is pretty consistent, and I thought it would be helpful to gather it here. Specifically, I've been thinking about ten-minute presentations, are often misunderstood. Below is some advice, and a video of an example talk.

The problem with short talks

There’s a lot of room for error in an hour-long talk. You have time to build a rapport with the audience, reacting to confused or sleeping faces by changing your pace or giving some extra background information. According to the peak-end principle, the audience will mostly remember the best part of the presentation and the end of it, so if you nail those you’ll be fine. People may also be there specifically to see you, so they start out primed to listen carefully and make an effort to understand. 

Not so for a ten-minute presentation: you get once chance to grab and keep the attention of an audience who might rather be playing with their phones.

In other words, you have to be better than Candy Crush.

If you only have ten minutes to speak, it’s probably because the organizers had to squeeze in too many talks into too short a time period. People will be tired of listening, will be thinking about their own talk, and will probably not have expertise in your field. So all of the decisions we make will revolve around three principles:
  • Give the audience an incentive to pay attention by being interesting or entertaining.
  • Don't make the audience work hard to follow your talk, or require prior specialized knowledge.
  • Make the most of your ten minutes by trimming anything unnecessary.
These principles apply in longer talks as well, but they're absolutely critical for short ones.

Here are some tips for putting these principles into practice:

  1. A 10-minute talk is not a shorter or faster version of a long talk. You can tell exactly one short story in ten minutes. Decide on the one thing you want your audience to take home with them, and write it down in a sentence. Don’t be afraid to repeat this sentence more than once in your talk. Your job is to provide just barely enough context to understand that story, and then tell it well. And no matter how slowly you think you're talking, you're talking faster than that. Slow down.
  2. The audience can either listen or read, but not both at the same time. Mostly you want them listening, so eliminate as much text as you can. If you’re afraid you’ll forget to say something, put it in the presenter notes. It can help to think of the slides as being there for the presenter’s benefit: they jog your memory so you can remember what comes next in the story.
  3. You will have the audience’s full attention when the talk starts. You will lose it immediately if you have a bad cover slide. Put the effort into making it tasteful but eye-catching, and spend time talking with that slide shown. That will buy you an extra minute of attention while you set up your story. Also, shorter titles are better. They're easy to understand, and the audience can listen to you instead of reading it.
  4. Don't use an outline slide. Outlines are useful when you’re going to talk about several topics and you don’t want your audience to get lost. In a short talk, you won't need to organize a complicated story, so outlines just eat the time you could be using to tell a simple one.
  5. Include animation only when it helps your audience follow the message. Make text come up a line at a time when you don’t want people reading ahead. That goes for figures too. You don't want the audience trying to digest your slide instead of listening to you.
  6. If you ask your audience to read text, make it easy on them. Using no font smaller than size 30 will not only guarantee readability, but will force you to limit the total amount of text on each slide (this includes chart labels!).  If you’re still using Comic Sans… don't.

    6.1: Related note on equations: include them only if they make the talk easier to follow. Equations are compact, efficient sentences that can be read in English. Being compact, they're difficult to unpack in real time. Don’t make the audience work that hard. Help them understand the symbols, and make it crystal clear why they lead to better understanding of the material.
  7. Spend the time and effort to make beautiful figures, and emphasize them instead of text. Help your audience understand them by pointing, and literally telling them "look over here." If you have graphs, always identify the axes out loud and teach people how to read them. Point out the important features. As usual, we want people listening instead of parsing data.
  8. End by saying “That concludes my talk. Thank you for your attention.” Don’t read your acknowledgements out loud (it eats time and gives the audience a chance to forget what they were going to ask), and don’t ask for questions (only the moderator knows how much time there is for questions).
  9. Make extensive use of supplementary slides. Paste entire presentations into the supplemental section, and load it up with equations and text. It doesn’t have to be pretty. This is your security blanket. If someone asks a question you can’t answer easily, you want to be able to find the answer in your supplementary slides. You’ll look like a genius for having thought ahead.
  10. Practice, but not to memorize your lines. Practice so you can find out what works and what doesn't, what's clear and what's not, and what you can safely trim away from the talk. Find the clunky transitions or a graphs that take too long to explain. There's no reason for a ten-minute talk not to sound smooth, relaxed, and well-oiled. If you can't get it to that point, you're trying to say too much.

Example talk

I made this talk as an example of how to implement some of the above suggestions, using a recent Python/Twitter project as subject material. It's not perfect (there was no easy way to record a pointer, for example), but hopefully you'll find it useful anyway.




Sunday, July 19, 2015

Data mining Twitter [with code]

I recently applied to the Insight Data Science Fellowship, and was invited to do a short Skype interview. The interview includes a short demo, which is supposed to show them what kinds of data I work with, and highlight some of the skills I bring to the data science table. Since I mostly work with MATLAB, I wanted to do a mini-project emphasizing some more relevant skills.

I got the email on Thursday with the interview scheduled the following Monday morning, so I needed something I could do over the weekend. Data mining Twitter seemed like a good option, since I could do it in Python (highly relevant for data science), it’s “real” data (as opposed to experimental data, I guess), and it lends itself to varied analysis including statistics and language processing. I just had to pick something to track.

Tracking #NorthFire

Friday night, there was a wildfire near Los Angeles, in the Cajon Pass. The fire jumped the highway, and about 20 cars burned. The hashtag #Northfire started trending, and I tracked it for about 14 hours, creating a database of about 5500 tweets, taking up about 32 megabytes of space.

First I’ll show the results, and then I’ll go into detail about the analysis. I’m also going to include the code I used, in case it's helpful to anyone. I made this graphic summarizing the analysis:

Summary of analysis for #NorthFire tracking over 14 hours. Thanks to my talented friend and freelance graphic artist Bethany Beams for helping me with this. She looked at my first draft and gave me some tips that improved readability substantially.

What did I find?

There’s not too much that’s surprising here. People were using the word “fire” a lot to describe the fire. Popular tweets include comparisons to Armageddon and references to exploding trucks. Standard, IMO. Tweets became less frequent as the fire raged on and people went to bed, and then picked up again when people started waking up and reading the news. One interesting thing is the popularity of the word “drone.” It turns out that some hobbyists had flown some drones in to get a closer look at the fire, which prevented the helicopters from dropping water. That’s why it’s important not to have a hobby.

Details on collection and analysis

I followed this wonderful tutorial to collect the tweets and perform some of the analysis. Collecting tweets basically involves:
  1. Registering an app with Twitter, which gives you access to their API
  2. Using Python to log on with your authentication details
  3. Using a package called “Tweepy” to open a stream and filter for a particular hashtag
  4. Saving tweets to file in the right format

Anatomy of a Tweet

A tweet is an ugly object. If you want to know how the sausage is made, look here:



It’s a database entry with the text of the tweet, the time it was created, a list of everyone involved, and about 30 other things I didn’t care about. The saved file is in JSON format, which is convenient for data science. Funny story, Twitter automatically supplies the tweets in this format, but Tweepy reformats them, so you have to manually change them back. Thanks, Tweepy.

Counting Tweets and Retweets

The next step was to count the number of originals and retweets, and save a new data file containing only the originals. This was important for the language analysis I wanted to do: having 600 retweets would seriously throw off the statistics. To find the retweets, I just looked in the text of the tweet, where retweets always begin with “RT”.

Most Common Words

I then used the file with the original tweets to track the most common words and bigrams. First, the text of each tweet has to be tokenized, where we parse the string of text into words and symbols. It’s also prudent to ignore “stop words” like “the”  and “a,” and punctuation. Python has a natural language toolkit that makes all of this pretty easy to do. Again, I used this tutorial.

Most Retweeted

Finding the most-retweeted tweets (getting tired of typing “tweets”) is similarly straightforward. I found some code here, but basically it just looks up the number of retweets for each tweet, puts them in order, and prints a list. You can set the minimum number of retweets and the number it displays.

Tweet Frequency Chart

The final thing I wanted to do was track the tweet frequency as a function of time. Each tweet contains a timestamp that reports the year, month, day, hour, minute, and second. I converted that to seconds using very straightforward code adapted from this page, and then saved all of the timestamps to a text file. I used MATLAB’s “histcounts” function to make a histogram, and plotted the counts as a line using the area plot function. In Adobe Illustrator, I recolored the histogram using the gradient tool.

Code

Here is the code on Github. Everything is a separate file in the interest of coding time. I may go turn each of the files into a function at some point. The important things are:

listen_tweets.py: Stream tweets from Twitter, filtering for a certain string. You have to put in your authentication details, like the consumer key and secret. You get those when you register an app.

discard_RT.py: For each tweet, check if it's a retweet. If not, save it to a file. Count the number of original tweets and retweets.

count_frequencies.py: Tokenize the text from all tweets in a file and find the most common words or bigrams using the natural language toolkit.

retweet_stats.py: List the most common retweets in order.

get_timestamps.py: Convert the "created_at" value from each tweet into seconds, and store all of the values to a .txt file.


That’s it for now. I hope this was helpful. Next time I’ll talk a bit about the interview.

Tuesday, June 30, 2015

Advice from a panel

I recently attended a panel discussion of four data scientists who are former particle physicists:

Eric Flattum
Distinguished Member of Technical Staff at Verizon
Michigan State University, PhD Physics (D0 experiment)

Kathy Copic
Director of New Initiatives and Growth, Insight Data Science
University of Michigan, PhD Physics (CDF and ATLAS experiments)

Heather Gerberich
Data Scientist, On Point Technology, Inc.
Duke University, PhD Physics (CDF experiment)

John Mansour
Vice President Advanced Solutions Group, Nielsen
University of Rochester, PhD Physics (E706 and CDF experiments)


Here is some of their advice that I found useful.

  1. "Data science" is a very vague term, and it refers to many different kinds of jobs. For each opportunity, find out exactly what you would be doing, and choose carefully. Jobs like "fit this Gaussian" and "look up something in this Excel spreadsheet" are going to disappear soon. Once you get into a company, you'll have room to move around, doing a wider range of projects or taking on management responsibilities.
  2. Keep track of what you're doing daily as a student or postdoc, and always think about where you could use those skills. We forget sometimes that we know how to do useful things, and you're going to have to sell your skills.
  3. Put together a portfolio of small projects. Emphasis on small. These can't be Coursera or class projects, since everyone looking for data science jobs have done these, and you need to differentiate yourself. The point is to provide evidence that you can work on real data to solve real problems. Try to pick something that's important to you. A panelst gave an example of a man who made an app to find hiking trails within some distance of his house by crawling the web. 
  4. In industry, there is a stereotype of PhDs that makes employers nervous about hiring them: that they like to work alone on difficult projects for a long time. Show them that you are willing to collaborate, and compromise between perfect and fast.
  5. Physics PhDs have three skills that employers like. Finding these skills in one person is apparently difficult.
    • Knowledge of analytics: a working understanding of basic techniques in machine learning, for example. 
    • Software development: the basic ability to import, process, and display data. 
    • Problem analysis: knowing when the data don't make sense, or something went wrong. Being able to estimate the most promising methods.
Something that I took home from this discussion was that my bar for saying that I understand something (say, Python) is waaaaaay too high. If you are comfortable answering questions about it in an interview, it can go on your resume.

- b

Tuesday, May 26, 2015

Why logistic regressions?

The Island

On a mysterious island live two types of animals: yeebles and zooks, and they love to eat wild strawberries. A certain cave on the island is home to either a yeeble or a zook, and you'd like to know which it is. Unfortunately, it's very shy, and never comes out of its cave if you're nearby. You decide to find out by feeding it strawberries.

From previous experiments, you know that yeebles tend to eat around four pounds of strawberries in 24 hours if left alone, while zooks will eat about eight pounds. You put ten pounds of strawberries outside the cave, and come back a day later to find 6.3 pounds missing. Which kind of animal lives in the cave? How certain are you?

Classification

This is a classification problem. Given the result of a measurement, we want to figure out what class a member belongs to. In particular, this problem is binary, since the animal is either a yeeble or a zook, and can't be both. It's also univariate, since the only variable is the number of strawberries consumed in 24 hours.

Suppose our data look like this:
Fig. 1: Previous data for two classes of animal
Each circle on this plot represents one animal. The horizontal axis shows how much the animal ate, and the vertical axis shows whether it was a yeeble (labeled 0) or a zook (labeled 1). Note that there is no sharp transition between the two types of animals. Some yeebles eat more than some zooks and vice-versa, so a measurement above 6 lbs/day does not guarantee that it's a zook.

Our strategy will be to turn the previous data about strawberry consumption into a function that takes [lbs of strawberries] and converts it to [probability that the animal is a zook]. The way this is normally done in the machine learning community is through logistic regression. We assume that the probability of an animal being a zook based on the amount of strawberries eaten is of the form

$P\left(Y|S\right) = \frac{1}{1 + \exp\left(a - b*S\right)}$,

where $a$ and $b$ are constants, and $S$ is the amount of strawberries eaten, in lbs. A regression is used to choose $a$ and $b$ so that the curve matches the data optimally. Just like in a linear regression, we choose the parameters that minimize a cost function like the sum of square errors.

Why the logistic curve?

Why do we assume a logistic curve is the correct model for the probability? Clearly, a line wouldn't work. At minimum, we need something with a range of $\left(0,1\right)$, since $P\left(Y\right)$ can't be less than 0 or greater than 1. But lots of functions have this property, including the sigmoids, of which a logistic curve is one. Choosing the logistic curve implies some assumptions about how the data are generated. 

In particular, the logistic curve is the correct probability distribution if each class exhibits a normally distributed feature with equal variance and different center values.

Here's an illustration of that assumption. Suppose we make a histogram of lbs of strawberries eaten for both yeebles and zooks:
Fig. 2: Histogram of previous data from which we construct our model. This plot shows 10k data points, wheras the scatter plot shows only 100 for clarity.
They're Gaussian-shaped (normally distributed), they have the same standard deviation, and they have different center values*. Again, it's clear that some yeebles eat more than some zooks, and vice versa.

These distributions are represented by

$P(S|Y) \propto \exp\left(-(S - S_Y)^2/2\sigma^2\right)$, 
and
$P(S|Z) \propto \exp\left(-(S - S_Z)^2/2\sigma^2\right)$,

where $S_Y$ and $S_Z$ are the centers of the distributions and $\sigma$ is the standard deviation. We want to know the probability that an animal is a zook given a measurement $S'$. By Bayes' rule, this is

$P(Z|S') = P(Z)\frac{P(S'|Z)}{P(S'|Z) + P(S'|Y)}$,

where $P(Z)$ is the prior probability of finding a zook (independent of strawberry consumption), and $P(S'|Y),P(S'|Z)$ are called sampling distributions. They tell us the probability that the data would have been generated if each hypothesis were true. They're just the Gaussian distributions given above.

We're further going to assume that $P(Y)=P(Z)$, that the number of yeebles and zooks is the same. If we knew otherwise, we would modify this. I'll talk a lot more about priors in other blog posts. Substituting in the known distributions and reducing, we have

$P(Z|S') = \frac{1}{1 + \exp\left((S_Z^2 - S_Y^2 - 2S'\left(S_Y - S_Z\right))/2\sigma^2\right)}$.

This is a logistic curve. We have shown that under the assumption that the measured features are normally distributed with equal standard deviations, a logistic curve is the correct probability model to use. This analysis extends to multivariate distributions as well.

Mathematical convenience

Logistic curves have some convenient properties, so we're lucky everything turned out this way. The whole shape of the curve is controlled by the argument in the exponential, $(S_Z^2 - S_Y^2 - 2S'\left(S_Y - S_Z\right))/2\sigma^2$, which is linear in $S'$. That means that we can use a generalized linear regression to find the parameters $a$ and $b$ above, which in this case turn out to be $a = 18.25$ and $b = 3.058$.

We can also easily show (by integrating the log of the likelihood ratio) that the probability of making an incorrect classification decision in this case is given by

$P(\textrm{fail}|S') = \frac{1}{2}\textrm{erf}\left(\sqrt{S_Y^2 - S_Z^2}/2\sigma\right)$, where $\textrm{erf}$ is the error function (another sigmoid!). Knowing this lets us address questions like "how close can the distributions be before we can't classify very effectively?" You tell me what error rate is acceptable, and I'll tell you how far the distributions have to be separated to achieve that error rate.

The answer

If we observe that 6.3 lbs of strawberries are missing after a day, all we have to do is check what probability it corresponds to on our logistic curve. Here is what the best fit curve looks like:
Fig. 3: The best fit logistic curve for the data allows us to classify future members.
In this case, l the answer is P(Z|S' = 6.3) = 0.625, There is a 62.5% chance that the animal is a zook.

Mathematical inconvenience

I want to briefly break the problem here to illustrate the limitations of the logistic model. We assumed above that the standard deviation of each distribution was the same, but that's not always a valid assumption. In fact, it's quite common to find Gaussian-distributed processes (or approximately Gaussian) where the standard deviation is proportional to the center value. If the standard deviations are different, the argument in the exponent is quadratic in both $S_Y$ and $S_Z$, so a linear regression doesn't work any more.

In fact, we can keep adding higher order moments to the Gaussian distributions, and the argument becomes some kind of higher-order polynomial, and we can use a polynomial regression to do our classification. 

What if the distributions are not Gaussian distributed? Well, that sucks as always. The sampling distributions won't be as nice, and P(Z|S') may not be parameterizable as a polynomial at all. Then we would need to perform general nonlinear fitting. That can be done of course, but not as efficiently.

I leave this as an exercise to the reader. ;)

- b


*: Gaussians actually extend over the entire real line, and I'm truncating at 0 lbs since it's hard to eat a negative weight of strawberries. The error incurred is small if the center of the distributions is a few times greater than the width.

Acknowledgement: Thanks to Thomas Stearns for helpful discussions.

Thursday, May 21, 2015

A study plan

Now that I have a better idea of what I need to learn, it's time to make a study plan. The point is to put things in an order that makes sense, knock off the highest priorites as early as I can, and to set quantitative goals to meet.

I found a really lovely graphic (left) with an explanation of what a data scientist needs to know, and nature jobs has a fairly good article from 2013 on the subject. There are also good discussions on Quora. Taking these sources together, here are the major tasks I need to accomplish, in no particular order.


*) Get good at math, statistics, and machine learning
By "math," they mean algebra, calculus, and linear algebra. I may need to brush up on the last of these. A friend recommended this book, which is free, and has some overlab with machine learning.

*) Learn to code
Spefically, pick a first language. Python, I choose you! Learn it at Codecademy and Google Classroom. I've already gone through the Codecademy course.

*) Learn about databases
This is where they keep the data. I should at least learn SQL.

*) Learn about data munging, visualization, and reporting
Munging means putting the data into a digestable form, which I assume involves dimension reduction. I understand the principles, but need to look into it more. This category seems like it should come out organically when I start working on side projects for fun.

*) Start using Big Data
This happens after smaller projects succeed.

*) Get experience
Kaggle competitions, side projects, and the like. Has to happen once I'm comfy with the basics.

*) Internship, bootcamp, job
I'll apply for an internship, but I won't have a shot unless I make huge progress before then.

*) Engage with the community
I already read fivethirtyeight. I signed up for a couple of societies and followed some people on Twitter. That's a start. Not enough time in the day to consume the content created by popular data scientists.


Here's a plan that makes sense to me:

  1. Take the Coursera machine learning course. It's free, I have the required math chops, and I've already started it. There's about 43 hours worth of video in this course, plus assignments. They estimate 5-7 hours per week for an unknown number of weeks, so... that's not useful. Let's say three months, so I expect to be done by August.
  2. Simultaneous with (1), start programming. Python will be my default language, but I'll use something else if I have a good reason. I already have a project lined up, but it's in C++. I'll tell you about it soon.
  3. After the Machine Learning course, take more statistics. Intro to Statistics by Udacity, and/or OpenIntro Statistics.
  4. Simultaneously with (3), start entering Kaggle competitions and engaging the community there. Scale up to big data when possible.


During all of this, I'll keep up with the blog and try to post mini-projects. I think this is a good start. I have a quantitative deadline of finishing the ML course by August, and I have a project that must be done by the end of the summer that involves programming. Hopefully I have not succumbed to the planning fallacy.

- b

Monday, May 18, 2015

Bootcamps, fellowships, and headhunters

There's a large demand for data scientists right now, and many people are making similar career transitions to mine. A partial infrastructure exists to move people with advanced science degrees into data science by reorienting their skills a bit, and in this blog entry I want to describe the landscape as I understand it.
***

Problem statement:

A group of highly educated people with a strong quantitative background would like careers analyzing data, which requires a specialized skills set including programming and statistical analysis. Very little exists in the way of secondary or tertiary education in this exact skill set. Who can make money off this problem?

Solution:

Educator-headhunter hybrids. Headhunters are formally called recruiters. They freelance or work for independent firms who match job seekers with employers, and they always get paid by the employers iff the applicant accepts a job offer. Essentially, they make money by wading through the slush pile to pick out the best candidates so the employers don't need to spend time on it.

In data science, there is a large worker pool that is not well matched to fill a recruiting demand, but could be retrained in a short time, say 6-12 weeks. The hiring companies can't retrain the workers themselves because they don't have the expertise. It would be like hiring a management consultant who had to be retrained in management theory before starting. If they could train the workers, they wouldn't need them.

The response to this demand has been two kinds of trainer/recruiter hybrids, loosely called fellowships and bootcamps. Both require an application process. Both attempt to match graduates with employers, and both reap recruiting fees from successful employment matches. Here are the differences.

Fellowships:
  • Are free. Anyone who gets accepted gets a full ride.
  • Are pretty short in duration. Typically about 6 weeks with full time attendance.
  • Are rare. There are only two I know of.
  • Have a highly competitive application process and require a PhD for admission.
  • Aim at people who are around 80-90% of the way to becoming a data scientist
  • Have ~100% placement rate after graduation

Bootcamps:
  • Cost about $12-16k. Partial refunds are sometimes offered if graduates accept offers from their corporate partners.
  • Last longer, around 12 weeks.
  • Are plentiful and widespread (in major cities globally). I've seen about a dozen.
  • Are still competitive, but do not require a PhD.
  • Aim at people who are about 40% of the way to becoming a data scientist.
  • Have ~90% placement rate after graduation.
In addition, there are stand-alond courses that don't match people with employers.

Online programs:

  • Pretty cheap: ~$4k per program, or a few hundred dollars a month.
  • Can be part time, and last from a month to a few months.
  • Always available and open to anyone.
  • Available at many levels of prior expertise.
  • No job assistance
And then there's always

Doing it your damn self:
  • Free minus opportunity costs
  • Minimum 12 weeks, probably closer to a year.
  • Tons of resources, harder to find peers for support.
  • No application process.
  • Can go from 0-90% pretty easily I think. Last 10% would be hard.
  • Placement rate unknown, but it's not a full time commitment.
I still have to answer some questions about these options before I decide which to pursue.

How do people support themselves while attending these programs? The fellowships only admit PhDs, who are old enough that they can be expected to have families (fun facts: the average new physical sciences PhD is 30.1 years old [source]. A 30-year old has on average something like 1 child [not quite applicable source]).

Can I find a job where I can learn these skills as I go? Maybe an internship? It probably wouldn't cover all of the bases, but it might be a start.

What are my chances of getting a fellowship? Most likely pretty bad if I were to apply right now. Deadlines are coming up in a month, so I might be able to build an impressive portfolio before then if I dedicate myself to it.

What do the reviews look like? Are there obvious scams? Do they not really teach you anything you couldn't learn just as easily by yourself?

Are there other options I'm not aware of? Cheaper or faster programs at a college or university? Online certifications with similar placement rates?

Need to keep digging.

***

A comprehensive list of options like this can be found here.

- b

Sunday, May 17, 2015

Simple linear inversion [code]

A couple of days ago I was asked to make a figure showing a simple linear mixing/inversion problem. It's a very straightforward task, but I thought this would be a good opportunity to share some MATLAB code online for the first time.

You can find the code here. First I'll explain what it produces, and then I'll talk a little about the code itself.

My task was to show the basics of spectral mixing. In spectroscopy, light is used to distinguish different chemicals by which wavelengths they absorb or reflect. A spectrum looks something like this:

The horizontal axis is in wavenumbers, which is just the inverse of the wavelength, $k = 1/\lambda$. The vertical axis gives the percentage of light absorbed if shone through the test sample. If there's a peak, it means that the molecules in the sample can be excited by that wavelength, which is a clue as to what kind of chemical it is.

Here I'm just making up spectra: there are no chemicals that correspond to these Gaussian lineshapes, but I needed something I could do quickly, and Gaussians are quick. I wanted to create a sample with several chemical contributions drawn from a large library of chemicals. Each member of the library is assigned a Gaussian spectrum with a randomly generated height, center, and width. The library is called $\mathbf{A}$ and looks like this:
I randomly select a small number of these, say five. Then I have to decide how much of each chemical is in the simulated sample. I do this by choosing a coefficient for each spectrum randomly, and then normalizing to make sure the coefficients add up to 1. That is, I choose a percent contribution for each spectrum in the sample. If we represent this so-called density vector is by $D$, then the composite spectrum is just the product $S=\mathbf{A}D$. It looks like this:


The black line shows the composite spectru, and the colored lines show the individual spectra multiplied by their percent contributions.

So far, we've looked at the forward model. That is, you tell me what the sample is made of, and I can tell you what the measurement should look like. This model is really simple. It just says that you take a weighted sum of the individual spectra. It's a linear system, basically meaning that if you double the inputs you get double the output.

Now we want to do the inverse problem: you tell me what the measurement looks like, and I'll tell you what the sample is made of. In general, inverse problems are close to impossible. Imagine if I told you that I weighed an object and discovered that it weighed one pound, and I ask you to identify the object. You might be able to make a good guess, but you don't have enough information to solve the problem exactly. I'll talk a lot more about making decisions under uncertainty, but the point for now is that you can't expect all inverse problems to have unique answers.

This problem is different. Not only do we have enough information to solve it, but inverting linear problems can be done with one line of simple code. The density is recovered by the equation
\begin{equation}
D = \mathbf{A}^{-1}S
\end{equation}
where  $\mathbf{A}^{-1}$ is the inverse of $\mathbf{A}$. We could find it by hand, but MATLAB does it for us if we call the function pinv().

The result is a density for each member of the spectral library, most of which are zero since they're not in the sample. In fact, we can quickly check which members are non-zero (actually about a threshold, just in case of machine errors). We can compare the recovered members to the members we chose to make sure the inversion worked.

***
Notes on the code

The main code is spectral_mixing, and there are three supporting functions, create-library, plot_spectsum, and plot_composite_spectrum. These output nicer plots that MATLAB does natively.

If you want to export these graphics as .eps files, I recommend using the print2eps function included in the wonderful export_fig package by Yair Altman.