skip to main |
skip to sidebar
One of the indicators of the position of women in academic science is the proportion of women among academic staff in STEM departments. How fast can this indicator change?The first constraint is the number of vacancies available for women to be appointed to. For example, if a department has a turnover of 5% per year and the number of academic staff is growing at 2% per year then the overall vacancy rate is 7% per year.The next most important constraint is the proportion of women in the pool of potential applicants. If the pool of potential applicants is 50% women and no women leave the department then a department with a vacancy rate of 7% could achieve an increase in the proportion of women of 3.5 percentage points per year. In these circumstances a department could get from 25% to 50% women in 7-8 years. However, if turnover was around 3% and the number of academics was static then the best the department could achieve, if no women leave, is a growth in the proportion of women of 1.5 percentage points per year. At that rate it would take 17 years to get from 25% to 50%. Of course, this might well be an underestimate since it is unlikely that no women would leave over a seventeen year period.There is an additional complication. Suppose a department is able to make six appointments over a three year period and it makes the appointments from a pool that is 50% women. If recruitment is fair with respect to gender then the probability distribution for the number of women appointed will be a binomial distribution with N=6 and p=0.5. This gives a probability of 0.31 of appointing exactly three women, a probability of 0.34 of appointing two or fewer women and a probability of 0.34 of appointing four or more women (probabilities do not add to one due to rounding). Hence for time periods in which a small number of appointments are made fair recruitment processes could easily result in apparent growth rates between 2/3 and 4/3 times the expected rate.Conclusions:1. There are limits to how fast the proportion of women among academic staff in STEM can increase. Growth rates of a few percentage points per year are not unreasonable.2. Estimating the long-term growth rate from measurements made over short periods is futile.The second conclusion implies that simply collecting data on the proportion of women among new appointees is unlikely to reveal inadvertent bias in the recruitment process. While these data are necessary it is also necessary to assess the recruitment process against best practice established by large studies, such as that described in the US National Academies Report Gender Differences at Critical Transitions in the Careers of Science, Engineering and Mathematics Faculty. (A briefing on the report is available from the pages of the National Academies Committee on Women in Science, Engineering and Medicine.)
To a scientist this is an incomprehensible question. How can it not be important to interpret data correctly?However, if your intention is to motivate people to take action then it does not matter whether your interpretation of the data is correct so long as people believe it.A number of my posts over the six months have dealt with the interpretation of data. Why do I think it is so important to interpret data correctly?- Incorrect data interpretation allows critics to discount your arguments.
- Incorrect data interpretation allows critics to re-focus the discussion as an argument about what the data mean.
- Incorrect data interpretation leaves you running the risk that scarce resources are directed at solving non-existent problems while leaving actual problems unaddressed.
- The examples of poor data interpretation I have discussed include:
Do we really make the argument for more women in science by demonstrating such basic errors in analysis? 5. I have a pedantic streak.
This example of data that does not support the inferences drawn from it is quite subtle. In the Campaign for Science and Engineering in the UK Policy Document Number 8, May 2008, ‘Delivering Diversity: Making Science & Engineering Accessible to All’ is the graph shown in Figure 1. It is noted that more women entered the system as researchers and were promoted at every subsequent level. However, the comment is also made that, while in an equal world the graph would be flat so that the percentage of women at higher levels matched that at lower levels, there is no sign of the rate of increasing under-representation abating, i.e. the slope in 2005/06 is pretty much the same as the slope in 1995/96. But the graph shows that the percentage of women at each level in 2005/06 matches that of the next lowest level ten years previously, which is roughly what you would expect as each cohort moves through the system. Since the percentage of women among researchers rose for the slope to flatten the percentage of women among professors, senior lecturers and lecturers would have to have risen faster. This raises the question: is the fall-off in the percentage of women with grade the result of current discrimination or disadvantage preventing women from moving through the system from researcher to professor or is it the result of discrimination or disadvantage that occurred before 1995. If the latter it is hard to see what we can do about it now.
In fact the profile in 1995/96 can be evolved to the profile in 2005/06 without making the assumption that women are disadvantaged in appointments or promotions in that the number of women appointed as lecturers or promoted to senior lecturer or professor matches the number available in the pool from which they are appointed or promoted. The results of a model in which the percentage of women among researchers rises linearly from 24% to 30% and the percentage of women among appointments and promotions matches the percentage in the pool are shown in Figure 2. The results from the simulation for 2005 match the data rather well. (A detailed description of the model I used to produce the simulated 2005/06 figures is available at Google Docs.) This does not prove that women did not experience discrimination or disadvantage in appointments or promotion during the period 1995 to 2005. What it does show is that the data in Figure 1 do not, by themselves, show whether they did or they did not. Of course, discrimination or disadvantage must have occurred before 1995 in order to have a starting profile that was skewed to lower grades in the first place, but I was not aware that there was any dispute about that.In the US the National Academies Report: Gender Differences at Critical Transitions in the Careers of Science, Engineering, and Mathematics Faculty (National Academies Press, 2009) found that there were gender differences in hiring. Women were more likely to receive the first job offer than they were to be asked to interview and more likely to be interviewed but the proportion of women applying was smaller than the proportion among those receiving Ph.D. degrees. Women were also more likely to receive tenure but the proportion of women candidates for tenure was smaller than the proportion of women among assistant professors. Because the data are snapshot data it was not possible to distinguish between the possibility that women are more likely to leave before coming-up for tenure and the possibility that the numbers of women being hired as assistant professors had increased in recent years. There were no gender differences in promotion to full professor. Studies of this kind that examine the transition processes with large enough data sets to reveal differences are what is required to reveal whether and when women are disadvantaged.
In the previous post I used standard hypothesis testing techniques to show that interpreting data requires value judgements. The following are some musings set off by that exercise.This post also contains technical material about statistical inference. Thinking about statistical inference tends to cause pain in the brain so for those who don’t want to struggle through the technical stuff the important points are:- The statement ‘The result is not statistically significant at the 0.05 level’ does not imply that there is no effect, only that the data do not rule out that there is no effect. The corollary of ‘The result is not statistically significant at the 0.05 level’ is ‘We are still not that sure whether there is an effect or not.’ (‘Absence of evidence is not evidence of absence.’)
- Statements about statistical significance are statements about the data not about the effect. With enough data a tiny difference of no practical importance can be statistically significant and a practically important difference can be not statistically significant if there is only a small amount of data.
I had gained the impression from my undergraduate textbook on probability and statistics that hypothesis testing was a routine procedure based on ideas that had been around since at least the 1930s. This is true. It is also true that there is considerable controversy.As I understand it (and I’m not a statistician, so please comment if you know about these issues) one controversy goes back to a dispute between Fisher and Neyman, basically about whether it is necessary to consider Type II errors. In a significance test you calculate the probability, p, that you would have observed an effect of at least the size that you did observe on the assumption that there is no effect. If this probability is small you reject the hypothesis that there is no effect. If you decide to reject the ‘no-effect’ (null) hypothesis whenever p is less than 0.05 then in the long run the rate at which you will incorrectly reject the null hypothesis (Type I error) is 1 in 20. Alternatively, you can tell people the actual value of p that you obtained.
A statement of statistical significance is a statement about the data not about the hypothesis. It tells us how confident you are about rejecting the null hypothesis on the basis of the data you have obtained. In some circumstances, for example, a quality controller in a factory deciding whether to reject or accept a shipment of components, you would need to know not only the probability of being fooled by fluctuations into rejecting the ‘no effect’ hypothesis when it is true (rejecting a shipment when it is OK) but also of failing to reject the ‘no-effect’ hypothesis when there actually is an effect (Type II error) (accepting a shipment with an unacceptable number of defects). The Neyman-Pearson approach allows you to construct a test that, for a given probability of rejecting the null hypothesis when it is true, minimizes the probability of accepting it when it is false. (More accurately, it minimizes the probability of incorrectly rejecting an alternative hypothesis. The assumption is that if you are going to take action on the basis of the test then you will either reject the null hypothesis, implicitly accepting the alternative, or you will reject the alternative hypothesis, implicitly accepting the null hypothesis.) You can always make the probability of Type II error smaller by accepting a higher probability of Type I error. In real world applications there will usually be arguments about relative costs and benefits to inform the choice. Fisher took the view that this procedure was incorrect in science:
“It is important that the scientific worker introduces no cost functions for faulty decisions, as it is reasonable and necessary to do with an Acceptance Procedure. To do so would imply that the purposes to which new knowledge was to be put were known and capable of evaluation. If, however, scientific findings are communicated for the enlightenment of other free minds, they may be put sooner or later to the service of a number of purposes, of which we can know nothing.”
[Statistical Methods and Scientific Inference, 1956]
In this view you either reject the null hypothesis with some probability of error or you regard the null hypothesis as not (yet) proved wrong. This is basically the logical point that you can use data only to disprove a hypothesis not to prove it. It does not matter how much data you have that are consistent with a hypothesis, there is always the possibility that you will eventually encounter data that disprove it. In this case you can’t make the error of accepting a hypothesis when it is false because you would never accept any hypothesis. However, when people do need to choose what action to take based on data they act as though they accept a particular hypothesis.
One of the points of contention has been that people don’t just need to know whether there is an effect or not they also need to know how big it might be. This is important even if you have not been able to reject the null hypothesis. For example, if you are comparing the rate of death from heart attacks for patients undergoing treatment A with that for patients undergoing treatment B it would be quite important to know that the data that gave a high (greater than 0.05) probability of error for rejecting the null hypothesis gave a not much higher probability of error for rejecting the hypothesis that one rate was twice the other. In other words, how well does the test discriminate between the null hypothesis and other hypotheses. This information should be given as a confidence interval (the interval that you expect to include the true value of a parameter in 95% of cases) or a calculation of the power, which measures how well the test discriminates between different possibilities.Why do we tie ourselves up in logical knots calculating the probability that an effect of at least the size we have observed would have been observed if the hypothesis we are trying to show is false is true? Why don’t we just calculate the probability that, for example, two means are different, given the data we have observed? This is the heart of one of the great divides of twentieth century, and now twenty-first century, science. In the theory of classical statistical inference the probability that two means are different is not a meaningful concept. The two means are fixed; we just don’t know what they are. It is the estimates of the means that we calculate from the data that have probability distributions. There is an alternative method for statistical inference based on Bayes Theorem. There are two problems with this method. First, it requires us to interpret probability as meaning a degree of belief. This upsets many scientists because one person’s degree of belief in a statement might be different from another’s. They prefer an objective definition of probability. This is called the frequentist position because a popular interpretation of probability is as the long term frequency in a large number of trials, for example, if you keep tossing a fair coin long enough the ratio (Number of tails / Number of tosses) would be close to 1/2. Those who favour the degree of belief interpretation are known as Bayesians, for obvious reasons. The other problem with the Bayesian approach is that it involves using the data to calculate a final distribution for the parameter of interest, called the posterior distribution, from an initial assumed distribution, called the prior distribution. Unfortunately, there is no accepted method for choosing a prior distribution. This does not matter if you have lots of data because then the prior does not have much influence on the posterior. (If you are not familiar with statistical inference you may at this point wonder why people doing classical statistics do not need to assume a distribution in order to calculate p. The reason is that there is a wonderful mathematical result called the ‘Central Limit Theorem’ that implies that the sample mean must follow a normal distribution when the sample size is large.)In the context of using statistics on academic appointments in a university department to inform action, the weaknesses of the Bayesian approach could be regarded as strengths. People act because they believe something to be true not because they have failed to reject the hypothesis that it is false. Also, in most departments there will be a range of pre-existing views on whether recruitment is biased in favour of men – a few who are convinced it is, a few who are convinced it is fair, some who have no idea and possibly a small group who believe that women are favoured over men. Some of these views will be strongly held. Why not make these differences explicit by incorporating them in different priors? [If some of the above sounds vaguely like something you once learnt in a stats course try ‘The Cartoon Guide to Statistics’ by Larry Gonick and Woollcott Smith, Collins, 1993, ISBN-13: 978-0062731029.]
This post gets a bit technical so for number-phobes the point is that interpreting numerical data involves value judgements, specifically about whether it is better to be fooled by random fluctuations into devoting time and resources to fixing a process that is not broken or to risk not fixing a process that is broken.When scientists start thinking about the issues related to women in science their first instinct is to collect numerical data. Natural scientists tend to want to measure a baseline, make an intervention and then re-measure to see if the intervention made any difference. Social scientists tend to want to measure everything they can think of and then do a multi-variate statistical analysis. This is natural. It is what we have been trained to do. We also often hear 'We must have hard data, not just anecdote'.Indeed, we do need good data but we should be aware of the limitations of the statistical approach. Let's take the example of a university department that wants to check whether its recruitment procedures are unbiased. If it is a life sciences department it might expect that 50% of academic appointments would be women, so in this case we can define unbiased recruitment as meaning that 50% of appointments are women. Mathematically the problem of determining whether this recruitment process is unbiased is equivalent to determining whether a coin is equally likely to land heads or tails. Here are the results of tossing a 10p piece ten times: HTTTHHTTTT. Seven of the ten tosses turned up tails. Is my coin-tossing biased? Seven out of ten seems quite large. What happens if I try tossing the coin twenty times? Here is the result of that experiment: THTHHHHTTT HHHHHHHHHH, which is five tails out of twenty tosses. At this point I wondered whether there was something about the way I toss coins that biased against tails so I tried another twenty tosses with the following results: HTHTTTHTHT TTTHHTTTTT or fourteen out of twenty. So from ten tosses there were 70% tails, from the first twenty tosses 25% tails and from the second twenty tosses 70% tails. Overall there were 26 tails out of 50 tosses or 52%. How can we decide whether or not my coin tossing is biased when the results are so variable?Conventional (Fisher) hypothesis testing says that we should calculate the probability, p, of seeing an effect at least as large as that observed on the assumption that the null hypothesis, in this case, that my coin tossing is unbiased, is true. If p is less than some value, conventionally taken to be 0.05 (or 0.01 or 0.001), then the null hypothesis is rejected. This test reduces the risk of us being fooled by random fluctuations into thinking that my coin tossing is biased when it is not. Coin tossing follows a binomial distribution, shown in the graph for N=10 and N = 50, and the relevant probabilities are P(7 or more tails out of 10 tosses) = 0.17, P(5 or fewer tails out of 20 tosses) = 0.02, P(14 or more tails out of 20 tosses) = 0.06, P(26 or more tails out of 50 tosses) = 0.44. So, 70% of ten tosses is not statistically significant, 25% of twenty tosses is statistically significant, 70% of twenty tosses is not statistically significantly and overall 52% of fifty tosses is not statistically significant. Calculating statistical significance guards against what statisticians call a type I error, in this case, believing that my coin tossing is biased when it is not. There is another possible type of error, which statisticians call a type II error, which is concluding that my coin tossing is unbiased when, in fact, it is biased. This type of error is not particularly important in many contexts, for example, my coin-tossing. However, if we are talking about appointments in a university department we might be very concerned if it was concluded that the appointments process is unbiased when, in fact, it is biased. The Neyman-Pearson procedure divides the possible outcomes of a measurement into two regions: the rejection region where the null hypothesis is rejected and the acceptance regions where it is accepted. For example, for ten coin tosses we might reject the hypothesis that the coin is unbiased if we get two or fewer tails (P(2 or fewer tails in 10 tosses) = 0.055 or if we get eight or more tails (P(8 or more tails in 10 tosses) = 0.055. In this case the probability of rejecting the null hypothesis when it is true is 0.11. The probability of accepting the null hypothesis when it is actually false depends on what the true value of the parameter p of the binomial distribution is. In this case, the probability of accepting p=0.5 when is 0.62 when p is 0.3, 0.82 when p is 0.4, 0.89 when p is 0.5, 0.82 when p is 0.6 and 0.62 when p is 0.7. So we have a better than even chance of accepting that the coin-tossing is unbiased when in fact the probability of getting a tail on any toss is anywhere between 0.3 and 0.7. If we change the rejection criterion to three or fewer tails or seven or more tails then the probability of rejecting when the hypothesis is true is 0.34 and there is a better than even chance of accepting the unbiased hypothesis when the probability of getting a tail on any toss is anywhere between 0.4 and 0.6. So, when we try to interpret our ‘hard fact’ that seven out of ten coin tosses came up tails in terms of making an inference about whether the coin-tossing process is unbiased the ‘hard fact’ disappears into a morass of value judgements about whether we prefer a higher probability of rejecting the hypothesis that my coin-tossing is unbiased when it actually is unbiased or a higher probability of accepting the hypothesis that my coin-tossing is unbiased when it is actually biased. In the case of academic appointments the judgements become do we prefer a higher probability that we devote time and resources to fixing a process that is not broken or is it more important that we are confident that the appointments process is not biased? We could, of course, use a larger number of tosses or appointments. The problem with this approach for academic appointments is that while it takes a few minutes to toss a coin fifty times it would take a department with fifty academic staff four years to make ten new appointments and twenty to make fifty new appointments if turnover is 5%. This does not seem a recipe for quick identification and correction of problems.