25 August 2008

Testing Ideas 2

Finally, in comments on my cherry-picking article, I was invited to take a look at http://rankexploits.com/musings/2008/ipcc-central-tendency-of-2ccentury-still-rejected The commentator didn't really say what I should be learning from the link or what relevance he thought the page had to cherry-picking. I discussed some basics about making good tests in part 1 on testing ideas. Please check out both prior posts as I'll assume you know their contents. It's a little work, I know, but in order to have intelligent disagreement, or agreement, we have to know what the other is saying.

The title of that article is "IPCC Central Tendency of 2C/century: Still rejected". Certainly catchy. But let's look at the substance: Does the IPCC say there's a central tendency (and if so, what is it that is supposed to have that central tendency) of 2C / century? An immediate red flag for me is that the site does not provide a reference to where the IPCC said any such thing. Being better about doing my homework, I went to the IPCC WG1 report itself and looked in the summary for policy makers and the chapter (10) which discussed the global climate projections. I invite you all to do so as well. In my look for 'central tendency', I found no such term in either section. In my reading, which has not been exhaustive, of the two sections, I still found nothing that could be construed as a 'central tendency'. The projections were projections, not forecasts, and were made for a number of different scenarios (assumptions about greenhouse gas emissions and other things). Some included ensemble means, but none of those were 2.0 C. One was 1.8, but calling that 2 is back to the problem of deciding whether I'm tall and letting me round to the nearest foot or meter. Not even out of the title and already there are problems.

In the lead paragraph, the author writes "... compared to the IPCC AR4’s projected central tendency of 2C/century for the first few decades of this century." Again, I don't find the IPCC saying that, and again, the author doesn't say where the claim comes from. The nearest match I find is in the Summary for Policy Makers, p. 12, where it says "For the next two decades, a warming of about 0.2 C per decade is projected for a range of SRES emission scenarios." The author mutated a term of precision, two decades, in to a vagary, the first few. Then the fuzzy term 'about 0.2 C per decade' got cast as a hard term of 2C/century. If nothing else, in reading this site, you're not reading a reliable reporter. Once I've reached that point, I generally stop reading a source. There was no need to misrepresent the original. Nothing was saved or simplified.

In terms of part 1, remember, I mentioned that you had to be careful to make a good test, including that it had to test the thing at hand. In that case, you had to be looking at my height and weight in order to decide whether I was 'tall and thin'. In this case, you have to compare (in some 'good' way) observations that bear on the thing actually predicted by the IPCC. When you read chapter 10, you'll discover several things. One is that the variable being projected is the global average surface air temperature. Another is that the models have interannual variability (they'd better -- nature does, as I mentioned in the cherry-picking and detecting climate change articles). And you'll see that 30 years is the normal period for averaging temperatures to make climate conclusions; but even so, the projections were being (gingerly, you'll notice if you know how scientists write) made about 20 year averages. One final thing buried in those 100 pages: The projections assume that there are no volcanoes and the sun's input remains fixed at the observed average value. We haven't had any major volcanic eruptions since Pinatubo, but the sun has been quiet since 2001 (below the average output).

So several things to go in to making a good test:
  • We have to compare 20 year average vs. 20 year average (the variable being suggested as meaningful in the report, and 30 would be better).
  • We have to look only at global mean surface air temperature.
  • We have to adjust for the fact that the sun has been giving less energy than assumed in the projections.
  • In making the test, we have to allow for the fact that the 0.2 C/decade itself has error bars (due to interannual variability and between-model variability)
The author does none of these things, and presents no reason why they are not necessary. Instead, she computes trend lines for January 2001 - present, not 20 year averages. She includes satellite observations of temperatures through the lower-mid troposphere (the UAH and RSS), rather than using only surface air temperatures. She ignores that the sun has been quiet. And she makes no allowance for the error bars on the IPCC projection.

Even if absolutely everything that was done in the statistics were right, which I'm not in a position to say much about, the test is not a good test so the result is at best meaningless. Chances are good that it's misleading (tests that aren't good are usually misleading).

The satellite temperatures show a cherry pick themselves. But first, why they shouldn't be used in this context in the first place. The thing is, they're not observing the variable that is being predicted. While it is connected, which might lead us to thinking that it's ok to use them, it still may not be. Back to the question of whether I'm tall. I think that the data you want is a measurement from the floor to the top of my head. It's true that leg length is related to height, at least in the sense that tall people generally have longer legs. So maybe you'd accept that instead. And say you took Michael Phelps' leg length too. If you looked at the two, you'd conclude that I'm taller than Phelps. (His legs are short for his height, mine are long for mine.) You'd be wrong, however; the data don't address well enough the question you're asking.

The cherry pick is that only 2 of the 4 satellite temperatures were taken, and it happens that the two are the two which show the least warming (you'd have to know about this, which is easy enough to find if you look but isn't universal knowledge). The author gives no reason for this selection. Further, even ignoring that, one of the two also is not a global data set. The RSS goes only 70 S to 82.5 N. There are good reasons for this, and, conversely, to not be confident about the UAH figures, but they get technical. It suffices here that only 2 of 4 data sources were picked, one of them isn't even global, and, worse, neither measures the variable of concern.

Even without knowing anything about statistical methods, we can see that the given site does not support its headline (and have some question about the headline as well). All that is needed is to check what the source (IPCC in this case) actually said, versus what was being tested. They're different things, so the test doesn't tell us about IPCC's projections.

We can't make the converse conclusion from this, that the IPCC projections are correct. We didn't test that idea, so no conclusion about it is possible here. We only checked whether the site was making a good test. It wasn't.

A non-digression to something related. While I've taken probability and statistics courses, they weren't very deep (I felt). I was thinking about taking more, and asked a coworker -- a statistician -- about doing so. We talked some about what I'd studied (and remembered, this was a good 15 years after I'd had the classes). She concluded that there was no real point for me to take the courses. More important, she said, was to understand the system I was working with. The appropriate statistical methods would suggest themselves, or at least I'd be able to hand a well-constructed question to a statistician. But without understanding the system, there was no point in applying statistics.

Project: Pull down the 5 data sources given at that site and compare the last 20 years (7/2008 to 8/1988) average global temperature against the previous 20 years, 8/1968 to 7/1988. Is the more recent average greater or not? Can't do that for the satellites as they don't go back that far. But try 8/1979 to 7/1994 versus 8/1995 to 7/2008. Only 15 years versus 14 years, so not very informative about climate, but it's all the data we have. Won't be a good test, but the best (perhaps misleading) that can be done given those data. (Try changing the periods too.)

johnathansawyer: as you invited me to look at the source, did the above affect your opinion about it? Why or why not? As usual, substantive reasons (either way).

[Update 28 August 2008] See my next two notes what is climate - 2? and evaluating climate trend for more on what appropriate averaging periods look like, and what happens if you compare the average of the last 7.6 years to the previous 20 years' average, respectively.

[Update 29 August 2008] rankexploits made a lengthy commentary on things related to this post, though misrepresenting even my first paragraph. Comments on that post are disabled (I get an error message in response to my comment attempt), my response is #5317 in http://rankexploits.com/musings/2008/ipcc-central-tendency-of-2ccentury-still-rejected

Folks who are getting heated about my 'advocacy', or defense of IPCC, or whatever. Take a minute. Read what exactly I said. Saying that a particular test was not good is far from saying that there can be no such test (I give an example of one that would be better myself!). Nor, as I said directly, does it mean that the projection is good. Let's see what happens when a good test (which mine isn't, just better) is made. Until one is presented, a poor test is still not useful.

23 August 2008

Blogrolling

Some sites whose interests are close enough to mine that they are linking here. My apologies to the folks whose language I was guessing at if I guessed wrong. Please, someone who knows any of those languages, have a look and let me know the correct language. Also, if I'm missing places, please do send a comment.

atmoz.org/blog Climate and Weather Explained -- See especially the oblate spheroid edition of the simplest climate model
www.emretsson.net/ (Swedish)
scienceblogs.com/clock/ A Blog Around the Clock
tamino.wordpress.com/ Open Mind
koillinen.wordpress.com/ (Finnish?)
stigmikalsen.wordpress.com/(Norwegian?)
thingsbreak.wordpress.com/ The Way Things Break
bravenewclimate.com/
chriscolose.wordpress.com/ Climate Change
www.scruffydan.com/blog/
rationallythinkingoutloud.wordpress.com/
simondonner.blogspot.com/ Maribo
rabett.blogspot.com/ Rabett Run
scienceblogs.com/deltoid/ Deltoid

[Update 2 Sept 2008]: See also
http://initforthegold.blogspot.com/
http://jules-klimaat.blogspot.com/

22 August 2008

Testing Ideas 1

I was invited (challenged, whatever) to take a look at a site that proposed to have disproved the 'IPCC prediction', over in comments to my cherry picking note. Here's part 1 of the look I promised in my comment reply over there. If you haven't already, please do read the cherry picking note. (Not only for my tiny little ego boost from having more page views, but because I'll be assuming here that you understand what all I mean by the term and examples of it for climate.)

Testing ideas is one of the central processes for science. Coming up with ideas is awfully easy. Supporting them takes work. Strengthening them so that they stand up to all good tests is extremely hard. But, they do have to be good tests. Same as it's harder to come up with supported ideas than just an idea, it's harder to come up with a good test than just a 'test'.

'Good test' does not mean 'comes up with the result I like'. Maybe it does, maybe it doesn't. You're usually much better off to not have specified before hand how you want the test to come out. A good test is one that is aware of the system it is studying, knowledgeable about the idea that it is testing, and has been devised carefully enough to confirm or deny the idea while also giving an idea of how firmly it is supporting or denying. Before launching in to the climate case, let's look at something very much simpler.

You might have the conjecture (even more tentative than a hypothesis), after having read my mention that I'm a distance runner, that I'm tall and thin. How would we test that? How can we make it a good test? You could just ask me. But that isn't a good test. You have no idea what I think is tall, nor do you have any idea what I think is thin. So if I say 'yes', you still don't really know anything. Poor test. You could try asking me my height and weight. But then you've made a poor test because you didn't establish what constituted tall or thin. If you're biased about the conclusion, it is far too easy to say after the fact that whatever figures I gave you do, or don't, constitute 'tall and thin'.

So you have to be precise about declaring what you're testing, what constitutes a pass and what constitutes a fail. You might arrive at something like 'taller than 80% of men your age means tall and lighter than 80% of men your height is thin'. Someone else might want those figures to be 90%, but, since you've specified how you arrived at your labels (and, of course, you will be sharing your data), they can examine for themselves whether they agree with your conclusion. You're still not out of the woods, however, because you don't know how I'm measuring my height and weight. Perhaps I'm wearing thick socks, shoes, and standing on my toes. Maybe I'm weighing myself fully clothed and carrying a backpack. You need to specify the conditions of the measurement as well. Further, you'll have to tell me how to report it. I could be 5'7" (170 cm) and round to the nearest foot/meter to tell you that I'm about 6' (2 meters) tall. That sort of ambiguity can completely ruin your test.

In any case, even for something as simple as deciding whether someone is 'tall and thin', you see that there can be quite a lot involved. One last thing, which winds up often being important in looking at people's conclusions. After going through the work of making a good test, you can only draw your conclusion about the thing you were testing. Suppose you decided that I was indeed tall and thin. That was your test, so that conclusion is reasonably good. What you cannot do, however, is conclude that because I am tall and thin, I'm a good distance runner. You never tested that, and haven't presented evidence that all tall and thin people are good distance runners. (They're not, nor is it true that all short and not-thin people are poor distance runners. Now you have to go back to the drawing board and decide what 'good distance runner' means and how to measure it.)

If something that simple involves so much work, something as complex as climate probably takes quite a bit more care. So, irrespective of much else, one thing to look for in reports about what is or isn't the case about climate is whether the author showed as much care in making their tests as I suggested for something as trivial as deciding whether someone was 'tall and thin'. In part 2, finally, I'll get to the examination I was invited to make.

21 August 2008

Labelling instead of thinking

It's been interesting to see how rapidly some have labelled this blog and what labels are out there. Almost always, hence this note's title, labelling is done instead of thinking. Labels around the topic of climate include alarmist, believer, skeptic, denialist, and probably several others. Conspicuously absent is 'science-minded'. If you have to have a label for this blog, that'd be one to use. I like science and think it's a good way to answer scientific questions. Not all questions are scientific, but when they are, it's good.

So I'm puzzled that there's been reference to me being a 'believer'. What in, the source didn't say, and that's rather the problem with such labels. People using labels instead of reality often don't tell you what the label means. I do believe that science is a good way to answer scientific questions, but if that's all it takes to be a 'believer' then all scientific sources would have to be labelled so. Perhaps they are, which leaves me wondering what the utility of their classification is supposed to be. Help people avoid learning about the science?

Denialists, as I might use the term, used to be across the spectrum as to conclusions. That is, there were folks who insisted that sea level was going to rise and drown everyone (!) in the next few years (unless we all did what they wanted) and denied any and all evidence to the contrary. At the same time, there were folks who insisted that sea level couldn't possibly change except maybe to go down. And again, denied any and all evidence to the contrary. I don't see the former types much any more, but the latter have only gotten more vocal and seemingly numerous.

Skeptic ought to have been a good label for the science-minded. Instead, it's been appropriated by a very narrow viewpoint and, as usual for such swipings, one that is not skeptical at all. A true skeptic will sit down and look at all evidence regarding a point. They aren't foolish enough to think that there are only 2 sides. And they look even handedly at all evidence. They don't select only 1 'side' and make only that one defend their conclusion. They also use consistent standards of evidence. In my cherry-picking note, for instance, I mentioned some who (dishonestly) use trends from 1998 (or one or two other particular, even more recent, years) only to conclude that 'warming isn't happening'. If they were honest skeptics, they'd turn around and agree that warming is happening if 2009 were warmer than 1998 (or if, as already happened in one data set, 2005 were warmer). In practice, however, these fake 'skeptics' simply move on to some other point, never updating their conclusions in light of new information.

I'm also surprised to hear myself called 'alarmist'. Even less useful that 'believer', as folks throwing that label never do say what was alarming. Apparently they're scared to hear that there really is such a thing as the greenhouse effect. But I don't see that their being easily scared by reality should mean much to the rest of us. Again, the label has been taken by one particular group (the self-described 'skeptics', who also dishonestly took that label for themselves) to label only one viewpoint. Often, these same people turn around and make alarming statements themselves -- about how there'll be a worldwide depression if anything were done in response to climate change. That strikes me as much more alarming than a lab-tested statement about there being greenhouse gases.

In fairness, I do notice that I'm using a label myself -- 'unreliable'. On the other hand, I only use it after reading the source myself (and encouraging you to do so yourselves) and laying out exactly why I'm using it. It's a rather narrow usage, even more so since I'm only using it for sources which can be shown in error without knowing much science or math.

I'll encourage people writing here to not use the 'alarmist' 'denialist' and other such labels. If you think someone or some source (me included) is wrong about something, go ahead and say so and present the solid evidence (see my link policy) that supports your point. (Just saying you think they are, or I am, wrong is not enough and will likely be rejected regardless of who you think is wrong.)

20 August 2008

Greenhouse Misnomer

It's a nuisance that greenhouses don't work by the greenhouse effect. Some seem to want to make it out to be some sort of catastrophe or indictment of science instead. But it is annoying nevertheless.

The greenhouse effect's existence was discovered by Fourier in 1827. Namely, the surface of the earth is too warm unless something else were going on -- as we found in our simplest climate model. Digressing a second, Fourier realized this even though it was more than 30 years before anyone knew that there were greenhouse gases. Tyndall published his experiments, that showed water vapor (H2O) and carbon dioxide (CO2) among others to be greenhouse gases in 1861. Conservation of energy is a very powerful law -- you can reach the conclusion even if you don't know how, exactly, it comes about.

In any case, the idea was that greenhouses let the sun's rays in and didn't let the air in the greenhouse radiate energy back out. The definition of 'greenhouse gas' became this -- that they were transparent to solar energy (completely or nearly so) but absorbed earthly radiation (in the infrared, rather than the visible and near-infrared of the sun).

Unfortunately for purists, greenhouses don't work that way even though 'greenhouse gases' do. The way greenhouses work is to physically trap air near the surface of the earth. The sun's rays heat the surface (by radiation) and the surface heats the air in the greenhouse (by conduction). Since the greenhouse is closed, the warm air can't rise very far and the greenhouse stays warm.

Ok, those are two word pictures. Both are nice sets of words. But ... can we do an experiment that will show us which one is right? We can, and it was published in 1909. The experiment built two different greenhouses (small, but enough to test the idea). In one, the walls were the usual glass, which is transparent to sunlight but tends to absorb 'earthlight'. In the other, the walls were halite (rock salt) which is transparent to both sunlight and 'earthlight'. If the infrared properties of the walls matter, then the two greenhouses will reach different temperatures. If it is only important that there be a wall, the halite-walled greenhouse will get just as warm as the glass-walled. (Give or take a little for how well the two conduct heat.) The result was that the halite made just as effective a greenhouse as the glass. (Wouldn't work so well as a practical matter. Why? But just as effective for what was tested.)

So greenhouses don't work by the greenhouse effect. Planetary atmospheres, on the other hand, do. It's a nuisance that the words don't really apply thoroughly. But this is English. It's not like this is the first time there has been a definition that didn't make rigorous sense. We're talking in a language where flammable and inflammable mean the same thing after all.

Project: Build your own demonstration of the greenhouse effect versus greenhouses. See
http://www.wmconnolley.org.uk/sci/wood_rw.1909.html for the original experiment and a discussion of the significance.