August 21, 2007

The Coming Revolution in Scholarly Communications & Cyberinfrastructure

That's the title for the collection of essays in the most recent CTWatch Quarterly v3i3. There's an incredible array of essays on the future of scholarly publishing, all of them very interesting and worthwhile (I've not read all of the essays yet, but I will). Authors include such notables as Clifford Lynch, Paul Ginsparg, Timo Hannay, Stevan Harnad, Peter Suber and others. This is must-read stuff for everybody in science and libraries as changes to the way scholarship is published will affect virtually everything we do.

From the Introduction (which also has author bios) by Lee Dirks and Tony Hey of Microsoft:

By now, it is a well-observed fact that scholarly communication is in the midst of tremendous upheaval. That is as exciting to many as it is terrifying to others. What is less obvious is exactly what this dramatic change will mean for the academic world – specifically what influence it will have on the research community – and the advancement of science overall. In an effort to better grasp the trends and the potential impact in these areas, we’ve assembled an impressive constellation of top names in the field – as well as some new, important voices – and asked them to address the key issues for the future of scholarly communications resulting from the intersecting concepts of cyberinfrastructure, scientific research, and Open Access. All of the hallmarks of sea-change are apparent: attitudes are changing, roles are adjusting, business models are shifting – but, perhaps most significantly, individual and collective behaviors are very slow to evolve – far slower than expected.


And the complete TOC, with brief excerpts from the articles:

  • The Shape of the Scientific Article in The Developing Cyberinfrastructure by Clifford Lynch, Coalition for Networked Information (CNI)
    For the last few centuries, the primary vehicle for communicating and documenting results in most disciplines has been the scientific journal article, which has maintained a strikingly consistent and stable form and structure over a period of more than a hundred years now; for example, despite the much-discussed shift of scientific journals to digital form, virtually any article appearing in one of these journals would be comfortably familiar (as a literary genre) to a scientist from 1900. E-science represents a significant change, or extension, to the conduct and practice of science; this article speculates about how the character of the scientific article is likely to change to support these changes in scholarly work.

  • Next-Generation Implications of Open Access by Paul Ginsparg, Cornell University
    The technological transformation of scholarly communication infrastructure began in earnest by the mid-1990s. Its effects are ubiquitous in the daily activities of typical researchers, instructors and students, permitting discovery, access to, and reuse of material with an ease and rapidity difficult to anticipate as little as a decade ago.

  • Web 2.0 in Science by Timo Hannay, Nature Publishing
    Over the last 10 years or so, much of the discussion about the impact of the web on science – particularly among publishers – has been about the way in which it will change scientific journals. Sure enough, these have migrated online with huge commensurate improvements in accessibility and utility. For all but a very small number of widely read titles, the day of the print journal seems to be almost over. Yet to see this development as the major impact of the web on science would be extremely narrow-minded – equivalent to viewing the web primarily as an efficient PDF distribution network. Though it will take longer to have its full effect, the web’s major impact will be on the way that science itself is practiced.

    The barriers to full-scale adoption are not only (or even mainly) technical, but rather social and psychological. This makes the timings almost impossible to predict, but the long-term trends are already unmistakable: greater specialization in research, more immediate and open information-sharing, a reduction in the size of the ‘minimum publishable unit,’ productivity measures that look beyond journal publication records, a blurring of the boundaries between journals and databases, reinventions of the roles of publishers and editors, greater use of audio and video, more virtual meetings. And most important of all, arising from this gradual but inevitable embracement of technology, an increase in rate at which new discoveries are made and exploited for our benefit and that of the world we inhabit.

  • Reinventing Scholarly Communication for the Electronic Age by J. Lynn Fink and
    Philip E. Bourne, University of California, San Diego
    Cyberinfrastructure is integral to all aspects of conducting experimental research and distributing those results. However, it has yet to make a similar impact on the way we communicate that information. Peer-reviewed publications have long been the currency of scientific research as they are the fundamental unit through which scientists communicate with and evaluate each other. However, in striking contrast to the data, publications have yet to benefit from the opportunities offered by cyberinfrastructure. While the means of distributing publications have vastly improved, publishers have done little else to capitalize on the electronic medium. In particular, semantic information describing the content of these publications is sorely lacking, as is the integration of this information with data in public repositories. This is confounding considering that many basic tools for marking-up and integrating publication content in this manner already exist, such as a centralized literature database, relevant ontologies, and machine-readable document standards.

  • Interoperability for the Discovery, Use, and Re-Use of Units of Scholarly Communication by Herbert Van de Sompel, Los Alamos National Laboratory and Carl Lagoze, Cornell University
    One major challenge to the existing system is the change in the nature of the unit of scholarly communication. In the established scholarly communication system, the dominant communication units are journals and their contained articles. This established system generally fails to deal with other types of research results in the sciences and humanities, including datasets, simulations, software, dynamic knowledge representations, annotations, and aggregates thereof, all of which should be considered units of scholarly communication.

  • Incentivizing the Open Access Research Web: Publication-Archiving, Data-Archiving and Scientometrics by Tim Brody, University of Southampton, UK, et al.
    The research production cycle has three components: the conduct of the research itself (R), the data (D), and the peer-reviewed publication (P) of the findings. Open Access (OA) means free online access to the publications (P-OA), but OA can also be extended to the data (D-OA): the two hurdles for D-OA are that not all researchers want to make their data OA and that the online infrastructure for D-OA still needs additional functionality. In contrast, all researchers, without exception, do want to make their publications P-OA, and the online infrastructure for publication-archiving (a worldwide interoperable network of OAI -compliant Institutional Repositories [IRs]) already has all the requisite functionality for this.

  • The Law as Cyberinfrastructure by Brian Fitzgerald, and Kylie Pappalardo, Queensland University of Technology, Australia
    In the realm of collaborative endeavour through networked cyberinfrastructure we know the law is not too far away. But we also know that a paranoid obsession with it will cause inefficiency and stifle the true spirit of research. The key for the lawyers is to understand and implement a legal framework that can work with the power of the technology to disseminate knowledge in such a way that it does not seem a barrier. This is difficult in any universal sense but not totally impossible. In this article, we will show how the law is responding as a positive agent to facilitate the sharing of knowledge in the cyberinfrastructure world.

  • Cyberinfrastructure For Knowledge Sharing by John Wilbanks, Scientific Commons
    Despite new technology after new technology, the cost of discovering a drug keeps increasing, and the return on investment in life sciences (as measured by new drugs hitting the market for new diseases) keeps dropping. While the Web and email pervade pharmaceutical companies, the elusive goal remains “knowledge management:” finding some way to bring sanity to the sprawling mass of figures, emails, data sets, databases, slide shows, spreadsheets, and sequences that underpin advanced life sciences research. Bioinformatics, combinatorial drug discovery, systems biology, and an innumerable number of words ending with “-omics” have yet to relieve the skyrocketing costs and increase the percentage of success in clinical trials for new drug compounds.

  • Trends Favoring Open Access by Peter Suber, Earlham College
    While it’s clear that OA is here to stay, it’s just as clear that long-term success is a long-term project. The campaign consists of innumerable individual proposals, policies, projects, and people. If you’re reading this, you’re probably caught up in it, just as I am. If you’re caught up in it, you’re probably anxious about how individual initiatives or institutional deliberations will turn out. That’s good; anxiety fuels effort. But for a moment, stop making and answering arguments and look at the trends that will help or hurt us, and would continue to help or hurt us even if everyone stopped arguing. For a moment, step back from the foreground skirmishes and look at the larger background trends that are likely to continue and likely to change the landscape of scholarly communication.

(via STS-L and others)

August 17, 2007

Interview with Joel Spolsky

The current issue of the ACM magazine Queue (v5i5 July/August 2007) has an interesting interview with software development guru Joel Spolsky, author of the Joel on Software blog.

Some excerpts:


BF You’ve written a few books in the past few years, and I see many books piled up in your office. What are you reading now? Anything to recommend?

JS There aren’t many books on the state of the art in software development and software management right now, which are the kinds of things that I like to write and read about. Unfortunately, we haven’t moved beyond the anecdote phase, and attempts to move beyond the anecdote phase are usually just anecdotes with statistics.

Part of the problem is there really isn’t a science going on here, which is very frustrating. The same applies to an awful lot of business writing. It’s very easy to write a book called, for example, The Starbucks Principle or The Dell Way, and just bring up a whole bunch of random anecdotes and somehow tie them together thematically and pretend that this is a real thing that you can do and be successful at.

And, lo and behold, another moron then is going to try to attempt the same thing in his or her own company. It doesn’t work because it doesn’t apply or because it didn’t work in the original company, either—it’s just an anecdote that somebody pulled out of thin air. That’s one of the problems that this particular field has been suffering from.

What we do have are anecdotes from the elderly, people like me, or even the truly elderly, the Gerald Weinbergs of the world, writing brilliant things. Timothy Lister, Tom DeMarco, Ed Yourdon—these are people writing, “I’ve been here for a long time. I know I’m a curmudgeon, but let me tell you young folk what it’s like, da da da da da da da. Here are some examples. What this showed was da da da.” If you read a bunch of those anecdotes, you actually may learn something. That’s sort of the oral knowledge of our field.

*snip*

BF Since you’re passionate about how people interact with computers, I’d like to get your thoughts on the current buzz about Web 2.0 user interface technologies. Has this stuff made our lives better or worse?


JS It has made life a lot harder for programmers, but it has probably benefited users when the programmers do it well. When the user interface works in the way that it’s expected to, and the user model corresponds to the program model, then it’s great for users.

On the other hand, it’s much harder for programmers for several reasons. First, the number of languages you program in has increased. You’re always writing some JavaScript for the browser. And then you’re writing in whatever server-side application language you use—Java, PHP, etc. And then you’re writing in HTML and CSS. You’ve got a bunch of languages to learn, to know, and to keep track of.

Code lives in lots of different places. Just because you wrote a little function that calculates the time of day in the current time zone in one place—let’s say server-side PHP—does that mean it’s going to work for you in JavaScript? That requires all kinds of coordination. The number of things that can go wrong is absurd. You have to test on three or four different popular browsers, at the very least.

Once again, we have hit the world where you have to write really small amounts of code and make them very tight and very fast. For a while we had the luxury of being able to write really sloppy code and never bother optimizing anything, because computers worked “fast enough” and the amount of code didn’t really matter. When you’re writing client-side JavaScript, however, there’s download time, and then the code has to compile. There is a practical limit to the amount of code that you can write on the client side.

Friday Fun, Part Deux: First CD Edition

Keeping with the musical theme for today...

Over at Whatever, John Scalzi asks

Over at By The Way, I note today is the 25th anniversary of the commercial availability of the CD, and asked folks there, who are of an age to remember remember buying a first CD as something special, to recall what their first CD purchases were. I thought I'd ask it here, too.

So: What was the first CD you ever bought?

As I note over at Whatever, it was Tango: Zero Hour by Astor Piazzolla. I haven't listened to it in a while, so I'll have to give it a spin when I get home today. If you've never listened to Piazzolla before, you really owe it to yourself to give it a try. It's the most amazingly compelling music you can imagine.

Friday fun: Crossroads Edition

Eric Clapton has long been on of my favourite performers. Recently, he and a bunch of his blues rock buddies had a benefit concert (Wiki) for his Crossroads (Wiki) rehab facility in Antigua. A good chunk of the July 28, 2007 concert is available on the MSN site. DVD on the way, of course.

Check out the list of performers: Jeff Beck, Doyle Bramhall II, Eric Clapton, Sheryl Crow, Robert Cray, Vince Gill, Buddy Guy, BB King, Alison Krauss and Union Station featuring Jerry Douglas, Sonny Landreth, Albert Lee, Los Lobos, John Mayer, John McLaughlin, Willie Nelson, Robert Randolph and the Family Band, Robbie Robertson, Hubert Sumlin, The Derek Trucks Band, Jimmie Vaughan, Johnny Winter, Steve Winwood and Bill Murray.

August 16, 2007

Recently in the IEEE

Yesterday was the first day back at work in the library. The sabbatical is now over.

So, what am I reading this week? Some cool stuff from IT Professional, IEEE Annals in the History of Computing and IEEE Transaction on Education.

August 14, 2007

Smolin, Lee. The trouble with physics: The rise of string theory, the of a science, and what comes next. New York: Houghton Mifflin, 2006. 392

This one of those very rare books, books that make you truly smarter and more knowledgeable that when you started. What does Lee Smolin, physicist at Waterloo's Perimeter Institute, make us smarter and more knowledgeable about, you ask? First of all, the history of theoretical particle physics and the search for a theory that unifies classical physics and quantum theory. Second, the progress of String Theory as a unifying theory and it's alternatives. Third, the culture of the physics community and how it influences the first two.

It's also one of those books that just stops you in your tracks every once in a while. An insight or a story provoking intense reflection and concentration. I'd be sitting there, reading, and suddenly, staring off into space. It makes for a slow but worthwhile reading experience. Since the book treats a lot of fairly advanced topics in theoretical physics, it's also a pretty mind expanding experience, requiring a fair bit of comprehension to soak it all in. Any previous knowledge of string theory of other physics concepts will only enhance the enjoyment (and comprehension) of this book. My general physics knowledge is probably above average and there were a couple of parts where I struggled a bit. There were times when things seemed to make perfect sense while I was reading; then, after putting the book down for a bit, all comprehension simply vanished. On the other hand, as will become apparent later in the review, the hard-core physics stuff isn't really the main payoff of the book so if you find yourself skimming some of the particularly hairy parts in order to keep up your momentum, that's OK.

The main topic of the book has to do with the lack of really productive research in physics since the mid-1970s, when the Standard Model was set out. Since that time, the main focus of theoretical physics has been String Theory. However, as advanced as the theory is, there has been no experimental proof that it is valid. In this sense, compared to the insane pace of advances in the previous century (from atomic theory, to relativity to Quantum Theory to the Standard Model), physics is in a crisis. Smolin attempts to understand that crisis, both from a scientific viewpoint and from a more sociological/philosophical viewpoint as well. Now, there's no more hoary a cliche than the brilliant scientist that turns to philosophy of science in his dotage, mostly to his embarrassment, but Smolin is no geezer and he definitely doesn't embarrass himself is his attempt to understand why the incredibly bright community of physicists has failed to make significant progress in such a long time. Smolin is certainly not afraid to criticize the String Theory community for being too single-minded, for refusing to entertain alternative ideas about theoretical physics, or the physicists themselves for being a bit arrogant or dismissive of their colleagues.

This is a brave and worthwhile book. Read it to learn a lot of physics. Read it to learn a lot about the culture of physics. But definitely read it.

Two blogs to check out in relation to this book, one pro-String Theory and one more skeptical are the group blog Cosmic Variance (example here) and Not Even Wrong (here).

Here & there

A bit of catching up from around the blogosphere:


  • Video games: The invisible middleman of the industry by Ted Kritsonis. It seems that Canada is not only a world leader in video game production, but in software shops that build the tools the developers use to build their games.

  • Get a First Life! Reminds me of about a million science fiction stories about virtual worlds where people become so addicted that their physical selves wither away and die. There should be bibliography out there somewhere...

  • Biologists helping bookstores. It seems that people are moving ID, etc, books out of science sections in bookstores and into other sections. Hmmm. Sounds like a good idea. I should make sure we don't have any ID books classified as science in our collection.

  • The Seven Signs of Bogus Science. A bit of an older item, but still well worth reading. Pin it up on the notice board in your library, even.

  • Issues in Scholarly Communication blog. It's new to me, but it seems like a very worthwhile blog to watch.

  • Rebecca at Adventures in Applied Math has a great ongoing series which makes up a kind of Course on Supercomputing. It includes sections on parallel programming, Unix, batch scripts, makefiles and other topics. It's obviously not done yet, but I certainly look forward to more installments. Some of its a bit hairy, but well worth the effort to get the gist.

August 13, 2007

How to innovate like Google

Saying you're going to be innovative is an awful lot easier than actually creating an environment that truly encourages new ideas and is able to bring them to fruition. Probably by several orders of magnitude. One organization that seems pretty good at that process is Google. But how do they do it? What are the Google "rules of innovative organizations" that other organizations can at least hope to pattern themselves on?

Well, jump on over the Curious Cat Science & Engineering Blog and watch the YouTube video of Google's Marissa Mayer talking about the 9 ideas that encourage innovation.

A summary from Curious Cat:


  1. Ideas come from anywhere (engineers, customers, managers, executives, external companies - that Google acquires)
  2. Share everything you can (very open culture)
  3. You're Brilliant We’re Hiring
  4. A license to pursue dreams (Google 20% time)
  5. Innovation not instant perfection (iteration - experiment quickly and often)
  6. Data is apolitical (Data Based Decision Making)
  7. Creativity loves Constraints
  8. Users not money (Google focuses on providing users what they want and believe it will work out)
  9. Don’t kill projects morph them

August 10, 2007

Friday Fun: Body Art Edition

Take a jump over to The Loom and check out his carefully assembled collection of people who've had their love of science tattooed on their bodies.

August 9, 2007

Rosenberg, Scott. Dreaming in code: Two dozen programmers, three years, 4,732 bugs and one quest for transcendent software.

Full bibliographic info (title field not big enough):

Rosenberg, Scott. Dreaming in code: Two dozen programmers, three years, 4,732 bugs and one quest for transcendent software. New York: Crown, 2007. 400pp.

Every organization relies on software these days. Big custom systems, shrink wrapped commercial software, all the various protocols and programs keeping the Net running. Big organizations, small organizations, tech companies of course, libraries in particular are relying on the fruits of software developers mental labours more and more. And with the rise of Web 2.0 in libraries and educational institutions, our reliance on our programmers will only get more pronounced. But how much do we really understand about the art of software development and the strange and wonderful habits of programmers, systems analysts and all the rest of the software bestiary?

Not much, it seems. And that's where this fascinating insider account a a high-profile open source software project comes in. Salon.com co-founder and author Scott Rosenberg spent three years as a fly on the way on Mitch Kapor's project to create the ultimate Personal Information Manager (PIM), Chandler. Kapor's project was highly idealistic from the very beginning; the idea was that he would use some of his software-boom fortune to finance a project to make every one's lives easier: a PIM that is flexible, sharable and open, able to handle calendaring, email, note taking and events. Unfortunately, the project was also cursed with design difficulties and numerous delays, with a schedule that stretched out from one year to two and three years and beyond (and not even implemented today). The book includes a colourful cast of both obscure and well-known software luminaries (like Andy Hertzfeld), and goes beyond merely recounting the ups and downs of Chandler but also offers a kind of history of attempts to organize and systematize software development. Name-checking such great software engineering writers as Frederick Brooks, Rosenberg talks about the whys and wherefores of structured programming, object orientation and others. Many chapters mix details of the vagaries of the Chandler project with relevant discussions of theoretical topics in software engineering (such as trying to create truly reusable software modules) with more philosophical musings on the art of software development. Most of all, Rosenberg places us firmly inside the workings of a programming project from hell, complete with gory details, tales from the historical trenches and a bit of that fantastic theoretical discussion on why software is so hard. (So, what's it really like being stuck in the programming project from hell? Trust me, I've been there and this is a pretty good example of the real thing.)

There are a couple of really good bits that really stood out for me in this book, bits that resonated with my own experiences managing and developing software. On page 54 he has a discussion of death march projects and the optimism/pessimism dichotomy that all programmers live with and obsess with every day. Having done a couple of death marches characterized by such extremes, it really resonated with me. On page 75, he begins a discussion various programming languages and the almost religious zeal most programmers have for their favourite ones – I was a big fan of Fortran as a young programmer. On page 274, Rosenberg has a telling comment about programmers' historical blindness, their inability to learn from their mistakes, to use the literature to learn from other's mistakes. I like the way he puts it: "It's tempting to recommend these [pioneering software engineering] NATO reports be required reading for all programmers and their managers. But, as Joel Spolsky says, most programmers don't read much about their own discipline. That leaves them trapped in infinite loops of self-ignorance." I like to think that as a librarian collecting the literature of software engineering, I can help in a small way to make programmers more aware of their past.

On a lighter note, I also like the joke that Rosenberg puts on page 275-276:

A Software Engineer, a Hardware Engineer, and a Departmental Manager were on their way to a meeting in Switzerland. They were driving down a steep mountain road when suddenly the brakes on their car failed. The car careened almost out of control down the road, bouncing off the crash barriers until it miraculously ground to a halt scraping along the mountainside. The car's occupants, shaken but unhurt, now had a problem: They were stuck halfway down a mountain in a car with no brakes. What were they to do?

"I know," said the Departmental Manager. "Let's have a meeting, propose a Vision, formulate a Mission Statement, define some Goals, and by a process of Continuous Improvement find a solution to the Critical Problems, and we can be on our way."

"No, no," said the Hardware Engineer. "that will take far too long, and, besides, that method has never worked before. I've got my Swiss Army knife with me, and in no time at all I can strip down the car's braking system, isolate the fault, fix it, and we can be on our way."

"Well," said the Software Engineer, "before we do anything, I think we should push the car back up the road and see if it happens again."

I'm going to use this joke when I do IL sessions for CS and Engineering grad and undergrad students, and maybe even to break the ice at a departmental meeting.

A great book, an insider view of software development, a real insight into how programmers think and work and how software projects grow and evolve, sometimes how they careen out of control. So, who would I recommend this book for? A number of different constituencies would find this book useful and entertaining.

  • IT Managers would find this book very useful for its insights into the personalities of programmers as well as for its history of failed attempts to make a purely predictable engineering discipline out of programming.
  • Programmers would find this book terrific, seeing a lot of their own eccentricities in the many stories. As well, programmers would get a lot of insights into their pointy-haired bosses attempts to turn them into engineers rather than the free-spirited hacker-artists they see themselves as.
  • Families of the either of the two above groups will get valuable insight into the slightly deranged members of their families, their joys, obsessions and frustrations.
  • People that support or employ software developers or managers, such as scitech librarians, HR people in tech firms, venture capitalists in software firms. They will hopefully come to understand how and why software projects are created and sometimes crash and burn. Not to mention how to mentor and encourage developers to take advantage of what is known to improve productivity. The other books and articles listed in the notes are also a treasure trove of further exploration and information. I hate it when books like this don't have a proper bibliography – it makes it a lot more trouble to sift through the notes later on for further reading.
  • And really, anybody that uses software of any kind. And since basically everyone uses some sort of software these days, just about anyone would really appreciate this book. Understanding how the knowledge economy and the Internet boom is built from the ground up is certainly enlightening and important. You'll never see a bank machine, interact with a big company's insane internal systems procedures or even use a simple web application the same way. Understanding the challenges involved in getting these systems even close to right and the inevitability of their imperfections is an important revelation in the modern world.

August 8, 2007

Chapman, Matthew. Trials of the monkey: An accidental memoir. New York: Picador, 2001. 367pp.

Sometimes, you just get lucky with books. What with the Dover, Pennsylvania trial only a couple of years a ago and the opening of various creationist museums, the evolution vs. creationism controversy is never far out of the news. Of course, I've seen publicity for a lot of books about the issue and Matthew Chapman's 40 Days and 40 Nights: Darwin, Intelligent Design, God, OxyContin, and Other Oddities on Trial in Pennsylvania is one that I've been looking forward to reading. While waiting for the paperback of that one to come out, I was browsing at the discount table at the local bookstore when what should I encounter, but the hardcover Chapman's first book, Trials of the Monkey, about the original creationist trial spectacle, the Scopes Trial. For $7. My lucky day. So, I bought it. It hung around the house for a few months, as usual, and then one day I just picked it up and started reading the first few pages, on a whim. I had no plans to start reading it seriously as I was about 100 pages into another very interesting book. Best laid plans and all, I was hooked and raced through Chapman's fascinatingly complex and interrelated account of the twistings and turnings of his own life and the story of the Scopes trial.

The book really has three narrative threads going. First of all, Chapman's biography, the evolution of his life, a troubled and mixed up childhood through to odd jobs and finally as a successful Hollywood screenwriter and director. The first parallel thread is his plan to visit the site of the Scopes Trial (Dayton, TN) and attend the annual dramatic reenactment of the trial; this takes two parts, the first being a trip to Tennessee to research the trial and scout out the area and the second to attend the reenactment. The final thread that Chapman weaves into a couple of the middle chapters is the story of the Scopes Trial itself.

A few words about the sections dealing with Chapman's biography. Chapman has a gloriously checkered past. The great, great grandson of Charles Darwin, he first chronicles the devolution of his family line from the lofty heights of the great man to his own mediocrity (ie. What's more mediocre than Hollywood) via the alcoholism of his own mother. His childhood, adolescence and young adulthood were certainly ones of very little accomplishment and many brushes with authority as well as a few bizarre sexual obsessions and entanglements. His climb to a happy family life, albeit struggling with the pressures of Hollywood and his own hard drinking, is a happy way to end this particular thread.

His trips to Tennessee to take in the annual dramatic reenactment of the original trial takes up probably half the book. He visits with many creationists, interviewing them and following them around to gather information and get a feel for the ambiance of the place; the same with some of the local bigwigs. It's very interesting that he really makes no attempt to demonize any of them, almost going out of his way to present their best side as well as their lack of scientific rigour and their eccentricities. He often comes off as liking them personally, almost admiring their convictions. This thread ends rather strangely and I won't spoil the surprise. Needless to say, given his own biography, Chapman isn't able to end the book anywhere near the way he'd like.

The chapters where he describes the history of the Scopes trial begin with the story of George Rappleyea, the local businessman who dreamed up the whole mess as a way of promoting tourism and business for Dayton. It follows with a fairly straightforward description of the trial itself and it's aftermath. Well worth reading for a lot of the colourful interactions between the two camps, William Jennings Bryan and his merry band of creationists against Clarence Darrow and the evolutionists. While it was a bit more bare bones than I hoped, this section will lead me to the bibliography to find other works to fill in the blanks.

Which of the threads is the most interesting and compelling? Easily, Chapman's life story is the best part of the book, followed by the story of his visits to Dayton. I often found myself skipping ahead to the next relevant chapter to see what happens next in his various adventures. The story of the Scopes Trial itself is somewhat played down, not given the attention of the other two threads. But that seems appropriate in the context of the story Chapman wants to tell. He covers the details well enough, but just not with the elan of his more personal adventures.

Over all, this is a worthwhile and compelling story, filled with intimate and telling biographical detail and local Southern colour. Not strictly a science book, more of a cultural history mixed with biography. It didn't end up being what I expected at all from the book; I expected more straight reportage and less personal anecdote and cultural commentary. But on the whole, the unexpected combination worked well, being both entertaining and enlightening. If I ended up understanding a little less about the Scopes Trial than I'd hoped, I think I ended up understanding a little more about what makes Southern Christian fundamentalists tick.

I would recommend this book for any public library and any academic library that collects popular science. Of course, any library that has an interest in the evolution vs. creationism controversy can't do without this book. I look forward to reading his second book, the one about the Dover, PA trial, Forty days and Forty Nights.

August 7, 2007

And now we resume our regularly scheduled blogging

The family and I returned from our vacation yesterday, and a brilliant vacation it was. We spent nearly four weeks at a cottage near Ste-Agathe, Quebec, about 1.5 hours north of Montreal. The lake we were on is very small with pristine, clear water and doesn't allow power boats and only has about 25 or so cottages total. Activities included swimming, canoing, paddle boating, bonfires, reading, relaxing, watching old movies on VHS in the cottage, go karting, mini putting, bowling, only checking email once per week in the village, BBQing almost every night. We saw the most recent Harry Potter movie one day on a trip into Montreal and The Simpsons movie in Kingston on the way home. And speaking of Harry, we bought the new book in the village the day it came out; my wife and boys raced through it but I haven't gotten to it yet.

We spent a couple of nights in Kingston at the wonderful Casablanca B&B, which we would heartily recommend. We also went on the Kingston Haunted Walk, which was good cheesy fun.

My reading included:


  • Stolen by Kelley Armstrong
  • Clear and Present Danger by Tom Clancy
  • Days of Infamy by Harry Turtledove
  • Chinatown Death Cloud Peril by Paul Malmont
  • Dreaming in Code by Scott Rosenberg
  • Ghost Brigades by John Scalzi
  • Ghost Road Blues by Jonathan Maberry

Reviews of all the above are forthcoming both here and on the other blog; I wrote two and a half science book reviews while on vacation. I'm finding Smolin's Trouble with Physics a tough book to write about, for some reason. Both The Yiddish Policeman's Union by Michael Chabon and The Map That Changed the World: William Smith and the Birth of Modern Geology by Simon Winchester were started and are still in progress.

July 6, 2007

Summer blogging break

It's that time again. I'll see you all back here sometime around the first week of August. Have a fun, safe and restful summer.

I expect that I will be checking my email and the blog comment moderation queue about once per week.

I've fallen behind a bit on my book reviewing, with two finished books to write up while I'm off: The Trouble with Physics by Lee Smolin and the just finished Trials Of The Monkey: An Accidental Memoir by Matthew Chapman. I haven't bothered with a poll to help choose my summer reading this year, but I thought I would share the titles I've chosen to get me started:


  • The Chinatown Death Cloud Peril by Paul Malmont
  • Days Of Infamy by Harry Turtledove
  • Stolen by Kelley Armstrong
  • Clear and Present Danger by Tom Clancy
  • The Map that Changed the World: William Smith and the Birth of Modern Geology by Simon Winchester
  • Ambient Findability by Peter Morville
  • Dreaming in Code: Two Dozen Programmers, Three Years, 4,732 Bugs, and One Quest for Transcendent Software by Scott Rosenberg

You'll note that there are some science-y books among the bunch, a bit of a break with tradition for me. We'll see how that goes. In any case, I do expect to report on my summer reading, both here for the science books and the other blog for the novels. The reviews should make up the first bunch of posts when I get back.

A year of stats

I've been using Google Analytics to track my blog traffic for a little over a year now and so I've finally come to a point where I have 12 months of pretty good data – July 2006 to June 2007. This seems like a good moment to take a look back and see what's happened over the last year, especially since it more or less covers my sabbatical leave that started last August.

Before June 2006 I only used the extreme tracking service to monitor traffic. It seems to count things a bit differently from Google so I hesitate to do direct comparisons. However, during the year July 2005 to June 2006 I got a total of 8,727 hits, for a monthly average of 729. This past year, Google gives me a total of 26,928 page views for a monthly average of 2,244. As I'm writing this, I've already surpassed last year's total for July.

That's quite a dramatic increase; most of that is due to both an increased posting frequency (from one or two per week to 4 or 5 per week) as well as a concerted effort to post better. I really made an effort to do more than just quicky, newsy posts and concentrate of offering real commentary and analysis of important scitech library issues. As well, the My Job in 10 Years series proved to be quite popular as has the occasional interview series, raising the profile of the blog quite a bit. Needless to say, I'm very pleased with the increase. (And I would like thank all my visitors over the years for their time and attention.)

Blogging better has meant that I was mentioned more often in other blogs, which in turn meant more traffic. Interestingly, a majority of links from other blogs seems to have come from science blogs rather than liblogs. My niche, partaking of both the liblog world and the science blog world, is a small one but one that I'm quite happy with. Some of the major supporters of the blog out there include Walt Crawford, Coturnix, PersonaNonData, Jane, CuriousCat, TWiL and, of course, all the scitech library bloggers (whom I'm not going to attempt to enumerate, but you know who you are).

Another thing that made me want to do this post is Walt Crawford's post from a few weeks ago Getting your fifteen minutes where he talks about his own traffic (about 85K for May) and Steven M. Cohen's for the same period: about 1 million. Mine was about 3,500 for the same period. Reflecting on the differences I note that I'm not bothered by them; I certainly don't begrudge either of them their success. I do what I do, hoping to please myself and generate some traffic as a side effect. On the other hand, it did get me thinking about taking a closer look at my stats. An interesting note is that while my hits are quite a bit lower than Walt's, our current Technorati rankings are in the same ballpark: he's around 44K and I'm around 50K. The highest ranking liblogs (such as Librarian.net or Information Wants to be Free), Technorati-wise, are under 1K. Whatever that means, and I'm not sure Technorati tells us much that is useful. For what it's worth, Scintilla has me listed as the 123rd most popular science blog.

Below is a chart of the last twelve months of pageviews and unique visits. As you can see, there was quite a dramatic increase starting with January this year, right after the post on A&I Databases. It's more or less platformed the last couple of months. June would normally have been a slower than average month, what with all the conferences, but two 10 Years Series posts boosted the numbers.



So, here are top 10 lists from July 1, 2006 to June 30, 2007, with a bit of commentary on some of the interesting ones.


Top 10 Posts


  1. My Job in 10 Years: Collections: Further Thoughts on Abstracting & Indexing Databases. My most popular post ever by a significant margin. It's very gratifying as it's also one of the ones I'm most pleased with and that I worked hardest on.
  2. Best and worst science books. A post mostly pointing to various of John Horgan's science books list with some of my own commentary and lists. An odd post to make number two, but a lot of people seem to want to know about good and bad science books.
  3. Giving good presentations using PowerPoint. Another mixture of links to other blogs and my own commentary. A popular topic.
  4. The life of a CS grad student.
  5. My Job in 10 Years: Conclusion.
  6. Facebook is public not private. Most of the hits are from people trying to figure out how to view private details on Facebook. A bit creepy.
  7. My Job in 10 Years: Physical and Virtual Spaces.
  8. Friday Fun: Build your own Sherman tank. A hoot. A lot of people seem to want to build their own tank and my post linking to some instructional videos has struck a cord.
  9. Interview with Jane of See Jane Compute. I'm happy that this interview was so popular. Lots of the links were from either Jane's blog or Scientiae.
  10. An Interview with Alison Farmer. A mystery. Lots of people see to search on the name Alison Farmer. I'm not sure if they're looking for the one I mentioned or some other Alison Farmer.

A couple of honourable mentions: the tags for the 10 Years Series and my Computers in Libraries session summaries both got enough hits to make the top 10 but I decided to bump them in favour of real posts.


Top 10 Referrers

  1. Bloglines.
  2. LisNews. Including TWiL.
  3. Computational Complexity. Mostly trackback links from posts I've linked to. Lance Fortnow's final post is a huge referrer for me. The Web is a strange place.
  4. Scienceblogs.
  5. Google. I think this mostly refers various Google services like Reader & Gmail.
  6. Technorati. Other people checking up on who's linking to them.
  7. Libdex Library Weblogs.
  8. The Official Google Blog. Mostly links from trackbacks.
  9. See Jane Compute. A good number are from Jane's link to the interview.
  10. Curious Cat Science and Engineering Blog.

Interesting mix of referrers, especially the balance between library and scitech sources.


Top 10 Keywords

  1. Science librarian. I'm the number one hit on Google for this search! Unfortunately, at only 165 hits for the year, there's not that many people doing the search...
  2. Confessions of a Science Librarian
  3. Mamdouh Shoukri. The new president at York University. I did a post welcoming him when it was announced and it's attracted a fair number of hits. Now that Dr. Shoukri has actually started, It might generate a few more hits. (Oh, by the way, if you're reading this Dr. Shoukri, Welcome to York and good luck with your new job. I hope to show you around the Library in the fall!)
  4. Best science books.
  5. Alison Farmer.
  6. John Dupuis. People looking for me! Or one of the other John Dupuis's out there. I find these searches a little creepy. For what it's worth, I'm also the number one Google hit for my own name.
  7. Best science books 2006.
  8. Nerac. I did an interview with Mike Mahoney of Nerac and I think that attracts some hits.
  9. Librarian science. A variation on the theme.
  10. Confessions. The people that find me using this search, I always figure they're quite disappointed when they see the actual content ;-)

Some of these have various permutations and combinations (ie. Librarian sciences, confessions science librarian) lower ranked in the list. I haven't bothered combining any of them here, as I feel the raw list gives a good feel for what's going on. One day I may do a post on the strangest keywords.


Top 5 Book Reviews

I'm only going to do the top 5 here, as I haven't reviewed enough book over the last year to make a list of 10 meaningful. Note that the list is a strange amalgam of stats from this blog and the other blog, so take it with an even larger grain of salt than usual.

  1. Three Science writing anthologies. Reviews of the latest editions of the Years Best American Science Writing, Year's Best American Science and Nature Writing and the first science blogging anthology, The Open Laboratory.
  2. David Suzuki: The Autobiography.
  3. Balanced Libraries: Thoughts on Continuity and Change by Walt Crawford
  4. King of Infinite Space: by Siobhan Roberts.
  5. Republican War on Science by Chris Mooney.

July 3, 2007

Interview with Timo Hannay, Head of Web Publishing, Nature Publishing Group

Welcome to the most recent installment in my occasional series of interviews with people in the scitech world. This time around the subject is Timo Hannay, Head of Web Publishing at Nature Publishing Group, publishers of Nature and other associated journals as well as web products such as Connotea, Nature Reports, Nature Network, Scintilla, PostGenomic, Nature Precedings and others. Way back in May I was contacted by Natasha Ighodaro of Nature to see if I would be interested in interviewing someone to talk about some of their new web products. Eventually, she put me in touch with Timo. As it happened, Nature was in the middle of rolling out a bunch of web products, so it took a while to actually get the interview down on pixels. In any case, I'm very happy with the results and very grateful to Timo for submitting to such a long interview and for responding with such Candor. Enjoy!


Q0. Timo, please tell us a little about yourself, your background and how you ended up as Head of Web Publishing at Nature.

It’s quite a long story, so here’s a slightly abridged version: I’m a scientist with an undergraduate degree in biochemistry from Imperial College, London and an doctorate in neurophysiology from the University of Oxford. (My specialty was synaptic plasticity.) I finished my doctorate in 1994, followed by a year of postdoctoral research at Waseda University in Tokyo in 1994-5. Back in those days I was also a freelancer for The Economist, and through a colleague at their Tokyo office I got to know the people at Nature Japan too.

I’d always been a big fan of Nature — my dad bought me a subscription when I was about 18 (which I’ll admit is pretty geeky) and when I was at Oxford my first paper was published in the journal. When I met the people at Nature they were just launching Nature Medicine, and in my spare time I started covering medical research stories for them from Japan. I then lost touch for a bit when I went back to London to join McKinsey & Co.

I worked as a consultant in the UK and Japan for about three years, which was an intense and brilliant introduction to the world of business. But too many of the companies we were serving were in sectors that didn’t especially interest me. So, through a series of happy accidents, I ended up joining Nature’s Tokyo office, working full-time on business development. I had been into computers since I was a kid, and by then I was especially interested in the web. It also happened to be the case that doing more stuff online, and doing it better, was the biggest business develop opportunity for Nature in the Asia-Pacific region at that time. So that’s what I focused on: developing their Japanese website, and adding Chinese and Korean sites. In late 2000 I moved to Nature’s London office to work on the main site, Nature.com. As part of that move, Howard Ratner, Nature’s CTO, put me in charge of a new team (of about 3 people) called New Technology.

We experimented internally with things like RSS, SVG, RDF and other three-letter acronyms, but as a technical team the scope for us to turn these into new services or businesses was somewhat limited. Sometime around late 2004 or early 2005 Annette Thomas, Nature’s managing director and now my boss, decided to create a Web Publishing department with a remit to experiment with the web in a much more user-facing way. Since then the team has grown to something over 20 people. I love what I do because it’s at the intersection of my main interests: science, technology and business. I’m only sad that I don’t have much opportunity use my Japanese any more. ;-)


Q1. Some of Nature's recent journal publishing decisions have been quite controversial among librarians. Nature Physics is a good example. Do we really need another Physics journal?

I have very little to do directly with our journals because my focus is explicitly on non-traditional online products and services. So I can only give my personal option, which is that if there wasn’t a need for any of our new journals then people wouldn’t submit their papers or subscribe to them. I honestly believe that we do a much better job than most other scientific publishers, and that we create better products. That’s why they’re successful. The Nature Reviews series is a great example. Until they came out, the typical editorial and production standards for review journals were, in my opinion, very low. Nature Reviews set a new standard. Our other titles do the same in their respective fields, and considering how heavily read and impactful they are, they’re also extremely good value. Cynics may think that I’m only saying this because I work at Nature and they pay my salary, but in truth it’s the other way round: I choose to work here because I believe that Nature does great things (and I certainly didn’t move from management consulting to scholarly publishing in order to improve my bank balance ;-).


Q2. First Connotea, Nature Network, the Nature Blog, Second Nature (Nature in Second Life): you seem to be getting into Web 2.0/social software in a big way. What's Nature's strategy for these types of initiatives in the longer term?

I think it’s important to realise that we don’t just work on participative "Web 2.0"-type services. We also do a lot in the area of scientific databases (see http://www.nature.com/databases) and podcasts (http://www.nature.com/podcast). But to concentrate on the Web 2.0 stuff: we’re basically trying to identify ways in which scientists can use the web as a collaborative environment. The web isn’t just a broadcast channel or a convenient way to ship PDFs around , it’s a completely new kind of medium in which our "readers" can connect with each other.

Since our job is facilitating scientific communication, if we can’t help scientists to make the most of the web — the most powerful communication medium that humans have ever known — then we’re not merely missing opportunities, we’re simply not doing our job. So at Nature we’re trying a bunch of different things, often inspired by interesting ideas we see outside science (Connotea is clearly based on del.icio.us, and Nature Network on things like LinkedIn and Facebook), but always tailored to what we think will be of most use to professional researchers and clinicians. We’re trying to test the boundaries of what we can do, and we’re not afraid to fail, though of course we always do our best to succeed.

In line with many web-based companies, but in contrast to the scholarly journals business, our services typically launch in a fairly basic form, then we develop them in response to usage patterns and feedback. Now that we have quite a few different services, you can also expect to see them start connecting together more.

Second Nature is a bit different. We’ve been following Second Life for two or three years now, and I think it has the same kind of disruptive potential that the web had in the mid-90s. (Whether it realizes that potential depends on a lot of things, so it’s far from certain.) It could become a profoundly important medium for scientific communication and education, and we want to be there early, understanding its strengths and limitations, and working with early adopters among researchers and educators to find out how we can add value. So far it’s been a positive and eye-opening experience; I’m optimistic about the long-term prospects.


Q3. How has the uptake been for Connotea, Nature Network and Second Nature? Is there a critical mass yet to make these social environments compelling to scientists and others? How will these social networks tie into the core journal publishing business?

Connotea has a user base somewhere in the tens of thousands (the exact number depends on how active a person has to be to qualify as a "user"). Nature Network is much newer so is still in the thousands. I don’t know the visitor numbers for Second Nature, but in terms of active contributors I guess we have a couple of dozen people engaged in creating things on our virtual land, which now extends over three islands. Connotea has enough usage to create interesting second-order effects. For example, it can do quite a good job at recommending things to you based in what you’ve bookmarked. Nature Network activity has grown extremely rapidly in the 4 or 5 months since launch and is approaching a level at which we would expect see that virtuous circle in which usage (e.g., in the form of forum posts) drives yet more usage (e.g., other people coming in to read the posts). There are numerous ways in which these could tie into our journals — Connotea-generated lists of recommended reading, links to articles authored by people in your personal network, etc. — but we’re much more focused in making these services useful in their own right.

If we achieve that then we should also be able to turn them into successful standalone businesses, even though that usually won’t be through the traditional route of selling subscriptions. We’re also trying to open up these services for others to use. For example, Connotea has an API (application programming interface) that we’ve used to create tagging and "related article" functionality for the institutional repository software, EPrints. Other people have used it to do similar things with their own web and desktop applications. The Connotea code is also open-source, so there are quite a few private instances, for example behind institutional firewalls. Some of those people have also contributed code back to the open-source code, which is great because we can’t possibly develop all the requested features on our own.


Q4. And speaking of Web 2.0, peer review is a core value in science. There's a lot of experimentation going on out there with alternatives to peer review, even Nature has stuck it's toe into the water. Where do you think this is headed -- no big deal or long-awaited revolution?

My personal view is that peer review is headed for a revolution at some point, but the timing is extremely difficult to predict because it depends mainly not on technology but on various interdependent and imponderable social factors. It could be in a year or in twenty years. Having said that, there are many people at Nature who are much more knowledgeable than me about these things and who think we’re going to keep more or less the current model of peer review for the foreseeable future.

The reason I think they may be wrong is that I basically buy the "wisdom of crowds" argument: there are plenty of examples of the web causing new, open and collaborative approaches to replace traditional, closed and proprietary ones -- from open-source software to Wikipedia. You don’t always get a better result to begin with, which is why skeptics find it easy to be dismissive (as they were with both open-source software and Wikipedia in the early days). But as they evolve, and particularly as more people join in, they get better until the results match or even exceed the traditional approaches, often at much lower cost. (Anyone who’s read Clay Christensen's work will recognize this as an important part of his "innovator’s dilemma" argument).

I also believe that the web is particularly well suited to a "publish then filter" approach rather than the traditional "filter then publish" approach that was required when publishing was necessarily a physical-world process. As you can tell, my belief is based on rather abstract reasoning, and by looking for analogies outside science, so even I’m not completely convinced by it. But I’m convinced enough to know that we ought to be pushing the boundaries, because peer review is completely central to what we do, and if there’s a better way to do it then we ought to be the ones to find it. But at least in science, no one has found it yet.


Q5. Tell us a little about your new product Nature Reports, how it was developed and what need you see it filling in the scientific information marketplace. Who do you see as its main audience?

I can’t take any personal credit for the Nature Reports series, but I can tell you a bit about it. It consists of three sites — on Avian Flu, Climate Change and Stem Cells — that aim to serve a couple of purposes. First, they provide sources of scientific information on topics that affect us all, and that are all too often the subjects of spin or misinformation. We want to provide a place for non-experts to go where they know that the information is scientifically up-to-date and unbiased. They go into more depth than the mainstream media, but not so much that any interested and intelligent person can’t follow.

Secondly, we want the Nature Reports sites to become places where scientists, policy-makers, business people, and other interested parties can come together to learn from each other. Particularly in the three areas currently covered by Nature Reports, science does not and cannot operate in a vacuum. It must be willing to give and receive information in a way that will help us all — together -- to make wise decisions on questions that could affect our world for generations to come. For example, there’s no way that scientists can decide on their own what we should and shouldn’t be doing with stem cells, because those decisions are ultimately social and moral ones, but science needs to inform, and be informed by, the debate.


Q6. What do you think the future of print journal publishing is in 5 years? 10 years?

The vast majority of journals are already accessed mainly online. Many forward-thinking organizations are morphing their libraries into places to work and meet, not primarily places where documents are stored. Some libraries are even becoming entirely virtual. And that’s even before you consider the rise of scientific databases, which are just as important as journals in many fields and are all online. So I think we’re already in a world where scientific information is primarily digital. Within 10 years, I think most journals won’t any longer exist in paper form because there won’t be any point. (Though people will continue to print individual items for reading.) Nature and one or two other journals will be exceptions because they are effectively also magazines that contain news and commentary as well as research. Many people (including me) still prefer to read the "front half" of these publications in print, but eventually they too will migrate to e-readers of some sort. However, predicting the timing of that development has already caught out a lot of people who are much cleverer than me, so I won’t try here. ;-)


Q7. How about journal publishing itself? In 5 or 10 years will we be able to recognize whatever it is that journals have evolved into? Is the very nature of scientific publishing headed for some sort of transformation?

I think the concept of the scientific "paper" will remain intact (even if that name will seem increasingly anachronistic). There’s real value in this unit of publication, which tells a story by explaining how something previously unknown has become know through a particular set of experiments. But beyond that, there’s a lot of potential for change. Smaller units of discovery will be published -- whether through blogs or databases or whatever -- because the barriers to publishing them are now so low. This, in turn, will create the need for new services to find and collate this information, preferably in a personalized way, and new measures of scientific impact that take into account such contributions, which will be much smaller and more numerous than published papers.

Journals will become better linked, easier to search, and more dynamic. Many databases will take more seriously the need for curation, peer review, citability and archiving. In this way, journals and databases will be harder and harder to tell apart, and I think the distinction between them will ultimately become meaningless. In cases where journals don’t add much editorial value -- whether through filtering or otherwise improving the content -- the concept of the journal itself may start to erode as readers become ever more concerned with the paper they are reading rather than where it came from.


Q8. Can you tell us a little about Science Foo? It looks like a lot of fun -- not something we normally associate with science publishing.

Science Foo Camp is certainly one of the most fun and cool things I’ve ever done at work. It’s based on a meeting format invented about 5 years ago by O’Reilly Media, the influential technical book publisher run by Tim O’Reilly. They run an annual event for techno-geeks called Foo Camp. (“Foo” is a word computer programmers use to denote some arbitrary value or name -- like “x” in algebra -- but in this case also stands for "Friends of O’Reilly".) Basically, Tim and his colleagues invite 200-300 interesting people to their HQ in Sebastopol, CA for a weekend of self-organised demos, presentations, brainstorming, contraption-building and musical jamming (basically whatever people find interesting).

The great thing about it is the quality and variety of people there: software billionaires, technically precocious teenagers, engineers, scientists, writers — you name it. The only criteria are that O’Reilly consider them to be doing interesting stuff, and they want to introduce them to others. So it’s a bit like a giant, manic, weekend-long dinner party for geeks. They’ve become something of a legend in techno-land. Anyway, I attend a lot of O’Reilly conferences (it’s where I steal most of my best ideas ;-) and have known Tim for several years. Last spring, at his Emerging Technology Conference in San Diego, following a conversation he had had with Linda Stone (a brilliant ex-Apple and –Microsoft person with a keen interest in science and medicine), Tim suggested to me that we organize a Science Foo Camp. I thought it was a great idea.

We then spent a few weeks looking for a suitable venue, during which time Tim asked Eric Schmidt at Google, who loved the idea too. That was in late May or early June last year, and we decided to hold the event in August, so we had only two months to get lots of interesting scientific people to the Googleplex. We were really worried that it would be too short notice to get the kind of people we were seeking. We were also worried that scientists from diverse fields might not have as much to discuss with each other as people from the technical realm. We needn’t have been concerned: it was a great success. Attendees raved about it and several went away with not just new ideas but new collaborations. One thing that worked really well — aside from the great venue and format — is that we included some non-scientists in the mix. These ranged from technology people with a strong interest in science to sci-fi writers and others with cultural links to science but from outside research. I think those people helped to foster a truly interdisciplinary mindset. We’re doing the same again this year, though this time we have had a bit more time to plan it. I really hope that we’ll be able to do this every year from now on.


Q9. Who do you think your biggest competitor is? Open Access journals, other society or commercial publishers or even just the notion that everything is available for free on the web?

None of the above. ;-) To be honest, I don’t spend much time thinking about any of those. Open access will come about mainly through funder-mandated self-archiving, not author- or sponsor-funded journals. Of course we compete with other established publishers too, but they are a relatively known quantity. Your point about everything being free is related to an issue that I think is critical for publishers of all stripes: how to create viable business models that don’t involve charging for content (whether readers or authors). That’s not because I believe it’s necessarily going to become impossible to do charge readers, but it won’t always be the optimal (or even a viable) business model, especially for collaborative online services, so we need other options. In short, we need to get much better at monetizing traffic.

But to answer your question, I think our biggest competitor is the unknown grad student in his (or her) dorm room hatching a plan to turn scientific communication upside down in the same way that Napster, Google and Wikipedia disrupted other industries. Such people are a threat precisely because of their obscurity and lack of any historical baggage. You no longer need a lot of money, or even necessarily a strong brand, to succeed online. Good ideas and implementation are much more important. That drives almost everything we do. I hope that we’ll come up with the best ideas and implementations first, not mainly because of a commercial desire to out-compete others, but because that’s how we can best support scientific discovery.

I see myself less as a scientific publisher and more as a scientist who happens to work in publishing, helping information about ideas and discoveries make their way as quickly and efficiently as possible from their originators to those who can put them to use. If I ever thought I wasn’t being effective in that role, I’d find some other way to spend my time, probably outside publishing, but almost certainly connected with science.


Q10. Nature Precedings almost seems like the boldest of Nature's recent web offerings, nudging the larger scientific community into the same direction as, say, the physicists. What was the rationale behind introducing the service, and what do you see as it's place in the Nature suite of web products?

The basic rationale is that it’s in the interests of science for researchers to share their findings with each other as early and openly as possible. As you say, this already happens in physics through arXiv.org (and Paul Ginsparg, who runs that service, has very kindly offered his advice as we’ve been setting up Nature Precedings).

There are all sorts of theories about why it doesn’t happen so much in biology and other fields, but we thought the time was right to try and kick-start it. For one thing, there seems to be an increasing acceptance and understanding of the power and value of the web in enabling open collaboration, whether through domain-specific scientific databases or much more general services like Wikipedia. We were also able to get public support from some outstanding partners: the British Library, the European Bioinformatics Institute, Science Commons, and the Wellcome Trust (with more to come, I anticipate). This is key because the barriers to adoption are much more social than technical, and no one organisation has the right mix of skills and influence to pull this off on its own.

For the same reason, we’re also reaching out to other publishers. I expect a few of them will be cautious at first, but many of them clearly appreciate what we’re doing, which is about complementing the journal system, not competing with it, and about building an open federated system, not a closed proprietary one. For our own part, Nature Precedings helps us to engage with scientists at an earlier stage of the research process, which supports our traditional journal activities.

Also, by moving early we hope to be among the first to work out how best to make this kind of service economically self-sustaining. We’ve already made clear that that won’t involve charging for access -- and we’re working with some of our partners to set up open mirror sites to guarantee that.


Q11. Scintilla, PostGenomic, Nature Reports, even Connotea, all seem closely related to me, all about organizing information and bringing it all together. Are all these services coming together eventually or are they going to get more differentiated?

To be completely honest, that’s not yet certain because it depends on how people use those services, and what they tell us about their needs. But my expectation, and our current intention, is to steadily integrate them in a way that will allow information from one application to be used within another, and for users to hop between them seamlessly. Ultimately the distinction between these different services should therefore become less and less pronounced.

I think that’s a good thing because people just want help with their scientific information needs, they don’t want to have to work out whether Connotea or Scintilla (or whatever) is the answer, and they certainly don’t want to have to visit several different sites to conduct a single task. That doesn’t mean they will all turn into one monolithic application, but it should become easier for (say) Connotea users to access Scintilla functionality, and vice versa. We’ve certainly put a lot of thought into making that kind of integration possible, but to what extent we pursue it ultimately depends not on us but our users.

July 2, 2007

Weinberger, David. Everything is miscellaneous: The power of the new digital disorder. New York: Times Books, 2007. 277pp.

David Weinberger's Everything is Miscellaneous is one of 2007's big buzz books. You know, the book all the big pundits read and obsess over. Slightly older examples include books like Wikinomics or Everything Bad Is Good for You. People read them and mostly write glowing, fairly uncritical reviews. Like I said, Weinberger is the latest incarnation of the buzz book in the libraryish world. So, is the book as praiseworthy as the buzz would indicate or is it overrated? Well, both, actually. This is really and truly a thought provoking book, one that bursts with ideas on every page, a book I really found myself engaging and arguing with constantly, literally on every page many times. In that sense, it is a completely, wildly successful book: it got me thinking, and thinking deeply, about many of the most important issues in the profession, at times arguing every point on every page. On the other hand, there were times when it seemed a bit misguided and superficial in its coverage of the library world, almost gloatingly dismissive in a way.

So, I think I'll take a bit of a grumpy, devil's advocate point of view in this review. I am usually not shy pointing out flaws in the books I review, but this will probably be the first time I'm really giving what may seem to be a very negative review.

Before I get going, I should talk a little about what the book is actually about. Weinberger's main idea is that the new digital world has revolutionized the way that we are able to organize our stuff. In the physical world, physical stuff needs to be organized in an orderly, concrete way. Each thing in it's one, singular place. Now, however, digital stuff can be ordered in a million different ways. Each person can order their digital stuff anyway they want, and stuff can be placed in infinite different locations as needed. This paradigm shift is, according to Weinberger, a great thing because it's so much more useful to be able to find what we need if we're not limited in how we organize in in the physical world. In other words, our shelves are infinite and changeable rather than limited and static. Think del.icio.us rather than books on a bookstore shelf.

Weinberger is sort of the anti-Michael Gorman (or perhaps Gorman is the anti-Weinberger?) in that the former sees all change brought about by the "new digital disorder" as almost by definition a good thing. Whereas Gorman sees any challenge to older notions of publishing, authority and scholarship as heresy, with the heretics to be quickly burnt at the stake. Now, I'm not that fond of either extreme but I am generally much more sympathetic to Weinberger's position; the idea that we need to adjust to and take advantage of the change that is happening, to resist trying to bend it to our old-fogey conceptions and to go with the flow.

So, what are my complaints? I think I'm more or less going to take the book as it unfolds and make the internal debates I had with Weinberger external and see where that takes us. Hopefully, they're not all just a cranky old guy pining for the good old days but that we can all learn something from talking about some of the spots where I felt he could have used better explanations or substituted real comparisons for the setting up and demolishing of straw men.

The first thing that bothers me is when he compares bookstores to the Web/Amazon (starting p. 8). Bookstores are cripplingly limited because books can only be on one shelf at a time while Amazon can assign as many subjects as they need plus they have amazing data mining algorithms that drive their recommendation engines, feeding you stuff you might want to read based on what you've bought in the past and/or are looking at now. First of all, most bookstores these days have tables with selected books (based on subject, award winning, whatever) scattered all over the place, highlighting books that they think deserve (or publishers pay) to be singled out. On the other hand, who hasn't clicked on one of Amazon's subject links only to be overwhelmed by zillions of irrelevant items. It works both ways -- physical and miscellaneous are different; both have advantages and disadvantages. After all, the online booksellers only get about 20% of the total business, so people must find that there's a compelling reason to go to physical bookstores.

Starting on page 16, he begins a comparison of the Dewey decimal system libraries use to physically order their books with the subject approach Amazon and other online systems use. I find this comparison more than a bit misleading, almost to the point where I think Weinberger is setting up a straw man to be knocked down. Now, I'm not even a cataloguer and I know that Dewey is a classification system, a way to order books physically on shelves. It has abundant limitations (which Weinberger is more than happy to point out ad nauseum) but it mostly satisfies basic needs. One weakness is, of course, that it uses a hopelessly out of date subject classification system as a basis for ordering. Comparing it to the ability to tag and search in a system like Amazon or del.icio.us is, however, comparing apples to oranges. Those systems aren't really classification systems but subject analysis systems. The real comparison, to be fair, to compare apples to apples, should have been Amazon to the Library of Congress Subject Headings. While LCSH and the way it is implemented are far from perfect, I think that if you compare the use of subject headings in most OPACs to Amazon, you will definitely find that libraries don't fare as poorly as comparing Amazon to Dewey and card catalogues. And page 16 isn't the only place he get the Dewey/card catalogue out for a tussle. He goes after Dewey again starting on page 47; on 55-56 he talks as if the card catalogue is the ultimate in library systems; on 57 he refers to Dewey as a "law of physical geography;" on page 58 he again compares a classification system to subject analysis. And on page 60 he doesn't even seem to understand that even card catalogues are able to have subject catalogues. The constant apples/oranges comparison continued for a number of pages, with another outbreak on page 61-2, as he once again complains that Dewey can only represent an item in one place while digital can represent in many places; really the fact that Weinberger doesn't realize that libraries use subject headings as well as classification and that an item can have more than one subject heading, well I find that a bit embarrassing for him, especially at the length he does on about it. Really, David, we get it. Digital good, physical bad. Tagging good, Dewey bad. Amazon good, libraries & bookstores bad.

It was at this point that I thought to myself that in reality, even Amazon has a classification system like Dewey, in fact they probably have a lot of them. For example, the hard drives on their servers have file allocation tables which point to the physical location of their data files. At a higher level, their relational databases have primary keys which point to various data records. Even their warehouses have classification systems, as their databases must be able to locate items on physical shelves. Compare using a subject card catalogue to find books on WWII with being dropped in the middle of a Amazon warehouse! He sets up the card catalogue as a straw man and he just keeps knocking it down and it get tiresome that way he just keeps on taking easy shots.

Weinberger also misunderstands the way people use cookbooks (p.44). Sure, if people only used cookbooks as a way of slavishly copying recipes for making dinner, then, yeah, the web would put them out of business. But, people use cookbooks for a lot of reasons: to learn techniques, to get insight into a culture and way of life, to get a quick overview of a cuisine or a style of cooking, as a source basic information for improvising, to read for fun, to get a insight into the personality and style of a chef, to get an insight into another historical period. The richness of a good cookbook isn't limited by just recipes.

I have to admit that at this point I was tempted to abandon the book altogether, to brand it as all hype and no real substance, a hoax of a popular business book perpetrated on an unexpecting librarian audience. Fortunately, I didn't. There were more annoyances, but the book got a lot stronger as it went along, more insightful and more penetrating in it's analysis. However, I think I'll stay grumpy. (hehe.)

One of the more annoying arguments (p. 144) that I often encounter in techy sources is that the nature of learning and the evaluation of learning has changed so radically that we will no longer want to bother evaluating students on what they actually know and can do themselves, but rather will only test them on what they can do in teams or can use the web to find out. In other words, not testing without cell phones and the Internet at the ready. Now, I'm not one to say that we should only test students on memorized facts and regurgitated application of rote formulas; and I think you'd be hard-pressed to find many schools that only do that. From my experience, collaboration and group work, research and consultation are all encouraged at all levels of schooling and make up a significant part of most students' evaluation. Students have plenty of opportunity to prove they can work in teams and can find the information they need in a networked environment. But, I still think that it's important for students to actually know something themselves, without consultation, and to be able to perform important tasks themselves, without collaboration. Certainly, the level of knowledge and tasks will vary with the age/grade of the students and the course of study they are pursuing. If someone is to contribute to the social construction of knowledge they, well, need to already have something to contribute. In fact, if everyone always only relied on someone else to know something, then the pool of knowledge would dry up. The book asks some important questions: what is the nature of expertise, what is an expert, how do you become an expert, are these terms defined socially or individually, how is expert knowledge advanced, how is expert knowledge communicated? A scientist who pushes the frontiers of knowledge must actually know where they are to begin with. At some level, an engineer must be able to do engineering, not just facilitate team building exercises.

And little bits of innumeracy bug me too. On page 217 he's trying to make the point that the online arXiv has way more readers than the print Nature. ArXiv has "40,000 new papers every year read by 35,000 people" and "Nature has a circulation of 67,500 and claims 660,000 readers -- about 19 days of arxiv's readers." Comparing these two sets of numbers is a totally false comparison. What you really need to do is compare the total download figures for arXiv to the total download figures for Nature PLUS an estimate for the total paper readership. For arXiv does he think all 40K papers are read by each of the 35K readers for a potential 1.4 billion article reads? The true article readership is probably much, much smaller than that. As for the print, the most recent Nature (v744i7148) has 14 articles and letters; for a guestimate for a whole year print, multiply by 52 weeks and 660,000 readers equals a potential 480 million article reads; probably not everyone reads each article, but at least most probably at least glance at each article. For the print only. He doesn't even seem to realize that Nature, like virtually every scientific journal, has an online version with a potentially huge readership, which Weinberg in no way takes into account. It's clear to me that, at least based on the numbers he gives, what I can actually say about the comparison between the readerships for Nature and arXiv is limited but that they may not be too dissimilar. Not the point he wants to make, though. Again, the real numbers he should have dug up, but did not seem to want to use, was the total article downloads for each source.

Now, I'm not implying that print is a better format for science communication than online -- I've predicted in my My Job in 10 Years series that print will more or less disappear within the next 10 years -- but please, know what you're talking about when you explore these issues. Know the landscape, compare apples to apples.

I find it frustrating that in a book Weinberg dedicates "To the Librarians" he doesn't take a bit more time to find out what librarians actually do, how libraries work in the 2007 rather than 1950. (See p. 132 for some cheap shots) But in the end, I have to say it was worth reading. If I disagreed violently with something on virtually every page, well, at least it got me thinking; I also found many brilliant insights and much solid analysis. A good book demands a dialogue of it's readers, and this one certainly demanded that I sit up and pay attention and think deeply about my own ideas. This is an interesting, engaging, important book that explores some extremely timely information trends and ideas, one that I'm sure that I haven't done justice to in my grumpiness, one that at times I find myself willfully misunderstanding and misrepresenting (misunderestimating?). I fault myself for being unable to get past it's shortcomings in this review; I also fault myself for being unable to see the forest for the trees, for being overly annoyed at what are probably trivial straw men. Read this book for yourself.

(And apologies for what must be my longest, ramblingest, most disorganized, crankiest, least objective review. I'm sure there's an alternate aspect of the quantum multiverse where I've written a completely different review.)