Showing posts with label google. Show all posts
Showing posts with label google. Show all posts

October 29, 2008

The Google Books Search deal: A real game-changer

Take a gander over at the Official Google Blog for an announcement of the settlement of the court case between Google and various publishers over the Google Books Search service.

While we've made tremendous progress with Book Search, today we've announced an agreement with a broad class of authors and publishers and with our library partners that advances Larry's and Sergey's original dream in ways Google never could have done alone.

This agreement is truly groundbreaking in three ways. First, it will give readers digital access to millions of in-copyright books; second, it will create a new market for authors and publishers to sell their works; and third, it will further the efforts of our library partners to preserve and maintain their collections while making books more accessible to students, readers and academic researchers.

I encourage you to read the post as well as the text of an page on the Future of Google Books Search where the true game-changing nature of the deal becomes glaringly apparent. I'll quote the part on Accessing Books:
Accessing books

This agreement will create new options for reading entire books (which is, after all, what books are there for).

  • Online access

    Once this agreement has been approved, you'll be able to purchase full online access to millions of books. This means you can read an entire book from any Internet-connected computer, simply by logging in to your Book Search account, and it will remain on your electronic bookshelf, so you can come back and access it whenever you want in the future.

  • Library and university access

    We'll also be offering libraries, universities and other organizations the ability to purchase institutional subscriptions, which will give users access to the complete text of millions of titles while compensating authors and publishers for the service. Students and researchers will have access to an electronic library that combines the collections from many of the top universities across the country. Public and university libraries in the U.S. will also be able to offer terminals where readers can access the full text of millions of out-of-print books for free.


  • Buying or borrowing actual books

    Finally, if the book you want is available in a bookstore or nearby library, we'll continue to point you to those resources, as we've always done.


Wow. People will be able to buy online versions of books on GBS. Libraries will be able to license all the content on GBS. Millions of books in all disciplines and from all time periods.

I can't wait to see details on this, especially if there will be some sort of DRM, how printing will work, whether or not you'll be able to download to readers such as the Kindle. Of course, it will be really interesting to see what a site license for a large university will cost. Will it be the equivalent of our entire monograph budget? The implications and the choices that would imply are staggering. Talk about a rock and a hard place. This has the potential to completely transform the ebook business and the way libraries buy books. The traditional players in the ebook business will have to really focus on seriously adding value to their offerings, the way A&I services have to add more value in the face of Google Scholar. Libraries will be faced with a lot of choices, especially in the face of fears of putting all our eggs in one basket.

Of course, I also have to blow my own horn here a bit. Way back, almost exactly three years ago, when GBS was still called Google Print, this is what I wrote in one of the entries in the My Job in 10 Years series, with emphasis added:
It's already happening: the New York Times, Globe and Mail, Toronto Star, all the JSTOR journals, Google Print. In 10 years, these will be the hot commodities in our libraries, all the stuff that the students are so frustrated that they can't find online. Why not all the Canadian newspapers back to the first issue? Why not all the books in Google Print full text searchable (and readable, for a fee). Who doesn't want to license the full text version of Google Print when it's finished -- and it should have made some pretty good progress in 10 years.

Of course, GBS isn't finished, and in a sense will never be finished. We live in interesting times.

May 22, 2008

Google Scholar and the future of A&I databases

This is a bit of a "told you so" post inspired by something I saw on Open Access News the other day: Google indexes 90% of recent engineering research. The post mentions a recent article in the The Journal of Academic Librarianship, May 2008: John J. Meiera and Thomas W. Conkling, Google Scholar’s Coverage of the Engineering Literature: An Empirical Study. The abstract says it all:

Google Scholar’s coverage of the engineering literature is analyzed by comparing its contents with those of Compendex, the premier engineering database. Records retrieved from Compendex were searched in Google Scholar, and a decade by decade comparison was done from the 1950s through 2007. The results show that the percentage of records appearing in Google Scholar increased over time, approaching a 90 percent matching rate for materials published after 1990.


My first Google Scholar post was way back in 2004 and I think what I said then is just as valid today:
Winners & losers:

  • Loser: the A&I industry. Big time. Google Scholar is free, their products are definately not. Can they add enough value to the data they have to make it worth our (ie. libraries) while to subscribe to their services? No one's cancelling all those indexes this year, or even next, but what about five years from now? The key here is adding value. Google's product will be one-size-fits-all, always a bit overwhelming. Also, it will be probably be limiting itself to stuff online-only. Will Google get the metadata for journal backruns that aren't online and refer users to their local academic library?


  • Winner: students. Big time. Students want to use simple interfaces, easy searches with highly relevant results. If Google can deliver that with this product like with their regular search engine, this will be a hugely popular tool amongst students.


  • Loser: non-OA journals. More and more, if a journal's content is not online for free, it will not exist for the new generation of scholars. Why use journal A behind some weird pay-money-or-else screen when journal B has their articles right here. I know that you can get to A via your friendly neighbourhood proxy server/academic library, but really, at 3 am with the paper due tomorrow and the student doesn't even know where the library is on campus, that's not going to happen. Also, anyone not afiliated with a subscribing institution will automatically choose B. It's only a matter of time before Google puts a "Free full text only" check box on the screen. Open Access will mean survival for journals in the Google world. Not this year, not next year, but maybe in five or ten.


  • Winner: academic libraries & librarians. Yes. We're winners. Think of what this could do for our budgets! Finally we can demo tools in the classroom that the students will think are relevant! No more blank stares & sneers! But seriously, the advantages of basically using one interface are huge in terms of teaching students how to get the most out of their search experience. Google will continue to be overwhelming for many and confusing to some, so we will still have the role of helping students navigate. Oh yeah, we'll actually be able to spend more time on concepts like critical thinking, scholarly communication and all those information literacy standards we talk about but rarely have time to actually teach.


  • Loser: vendors of federated searching products. One search is here. This is it. The real challenge, of course, will be figuring out how to get link resolver products like SFX to work with Google Academic. Also, for us Ontario universities, all our content is on a central server. How do we get our students using Google Scholar to find the content on our platform rather than automatically going to the publisher's site. An interesting challenge.


  • Winner: the general public all over the world. Obviously, this will bring together a lot of information and make it accessible to everyone. As more and more stuff becomes OA, more and more scholarly content will become easily accessible to everyone. This is a good thing.


Musings on the future of A&I indexes also played a very important part of my My Job in 10 Years series, with a whole post devoted to the issue -- one of my all-time most read posts, if that has any meaning. I won't quote here, but my main point was that in a Google Scholar world, A&I providers will have to struggle to figure out how to add enough value to the bibliographic, citation and indexing data to make it worth our while as librarians to license those databases. The evidence from the study cited above would seem to indicate that we're getting closer to the day where we can start doing other things with that money. Sure, there's still quite a few cases where the vendors add tons of value to the data (SciFinder, Illustrata, Web of Science...), but for how much longer is it going to be worth the huge investment on our part. Personally, I'd much rather be spending the money on acquiring full text content, digitizing our own unique collections and new services to reach out to our patrons.

Some of the places I've talked about this (and related issues) before:

August 13, 2007

How to innovate like Google

Saying you're going to be innovative is an awful lot easier than actually creating an environment that truly encourages new ideas and is able to bring them to fruition. Probably by several orders of magnitude. One organization that seems pretty good at that process is Google. But how do they do it? What are the Google "rules of innovative organizations" that other organizations can at least hope to pattern themselves on?

Well, jump on over the Curious Cat Science & Engineering Blog and watch the YouTube video of Google's Marissa Mayer talking about the 9 ideas that encourage innovation.

A summary from Curious Cat:


  1. Ideas come from anywhere (engineers, customers, managers, executives, external companies - that Google acquires)
  2. Share everything you can (very open culture)
  3. You're Brilliant We’re Hiring
  4. A license to pursue dreams (Google 20% time)
  5. Innovation not instant perfection (iteration - experiment quickly and often)
  6. Data is apolitical (Data Based Decision Making)
  7. Creativity loves Constraints
  8. Users not money (Google focuses on providing users what they want and believe it will work out)
  9. Don’t kill projects morph them

June 18, 2007

WILU2007: Google and Beyond: What Sources are Students Really Using?

(Reposted from here.)

By Don MacMillan (University of Calgary)

Abstract

This was an interesting presentation about an IL evaluation project at the University of Calgary. They wanted to see what bibliographic search engines (free or fee) that students were actually using by their 3rd or 4th year. The project had some interesting and even surprising results.


The situation is one where there is long-standing integration of IL skills training in the biological sciences curriculum, including ongoing assessment. The goals of the the project was to promote student reflection on research skills and the changes in habits over time but mostly to see if IL instructional content was aligned with student needs. Hopefully the info could be use to market further IL to faculty. The subjects were 25 3rd & 4th year bio students, most with at least one IL session in the past. The survey tool used was the FAST tool: Free Assessment Summary Tool (http://www.getfast.ca). The tool anonymously summarized a 16 question survey of student impressions at the end of an IL session.

Top resources used by students, in order with percentages: pubmed (84%), bioabs (80%), library catalogue (76%), google/scholar (72%), web of science (20%) and patent search (8%). Some resources tended to have different uses with students: bioabs for exploring a topic, google for choosing a topic, pubmed for exploring a topic and finding specific info.

Some impressions from students:


  • On which resource to use first: Pubmed popular but Bioabs and Google Scholar coming up
  • What source is the most useful: Pubmed is most mentioned
  • How has research changed during studies: library sessions mentioned, as well as using a greater variety of search tools
  • What caused the change: instruction & tutorials as well as profs & TAs
  • What do you wish you'd known earlier: some details on how to use Pubmed
  • How did you learn about the various resources: mostly mentioned library sessions
  • Do you plan to pursue an advanced degree: 14 yes, 10 no

Unanticipated results: lots of positive responses regarding IL, high use of Pubmed, not much Google Scholar use, students recognize the value of research skills in finding a job.

Conclusions: students use a variety of strategies, reinforcement of skills throughout program works, students use different tools for different purposes, survey benefited students and librarian.

Future directions: incorporate Pubmed earlier, split Google & Google Scholar next survey, add Scopus next time, ask follow up question later in term.

Downloads: ppt, survey questions, results, related readings, IL web page.

April 13, 2007

Vise, David A. and Mark Malseed. The Google story. New York: Delta, 2006. Updated Edition. 326pp.

Ah, Google. The 800 pound gorilla. The elephant in the room. The bull in the china shop. Really, the kings of the online world. And to think, just a few short years ago, nobody had ever heard of them. Myself, I remember starting to use Google in 1999 or so, when the buzz around library school was this cool new search engine that had way better relevance ranking and a sparse, clean design. By 2000 or 2001 I remember thinking to myself that their product was so impressive that they must have been a huge, thousand employee megacorporation. Little did I know that for most of those early years, Google was still a small, intimate, human-scaled company that very much reflected its founders, Stanford PhD students Sergey Brin and Larry Page.

Google has been around so long, at least in Internet time, and has been so prominent and omnipresent in the media with so much detailed reporting on blogs and in newspapers and magazines, that I tend to think I know the story. But do I? Are there things that I don't know about the giant? As it turns out, yes, there were a lot of things I didn't know about Google and Vise & Malseed's book does a great job of filling in the blanks. And a lot of aspects of Google's story, both bits I knew and bits I didn't, have significant lessons for the library world.

Of course, this is really a business book, not a tech book or a history of science book or even a library science book, so how did I end up with it on my sabbatical science book reading list? I remember when the book came out in hardcover there was quite a bit of press and I thought I would probably want to read it eventually, it and The search : the inside story of how Google and its rivals changed everything by John Battelle. Although the business library ordered it, I never got around to checking it out, figuring I would just buy the paperback when it came out. So, it comes out in paperback last fall, but since I never check the business section of the book store I never noticed. A few weeks ago, while we were at the airport waiting for our flight to New York for our March Break trip, I was browsing at the airport bookstore. Now, airport bookstores are pretty small; they also cater to business travelers more than regular folk so the pb version of The Google Story was fairly prominently displayed at the front of the store. And I bought it and read most of it during the trip. Which makes me wonder, doesn't classification sometimes make it harder rather than easier to find something? And sometimes, a small, focused collection can lead to more serendipity that a big huge comprehensive collection. See, even how I found the book and ended up reading it have a lesson.

But, enough of the chatter. How's the book itself? Is it worth reading? Like I said, I would like to concentrate on the parts of the Google story that were interesting or new to me and how I think those apply to the Library world.

The first thing that really struck me (in chapter 3) was that in the early years a number of companies had a chance to license Google technology and passed it up. Altavista wanted a home-grown search solution while Yahoo! wanted people to stay on their own site rather than searching and leaving. The didn't realize how important search was, so they missed a golden opportunity. Chapter 8 goes into that idea in more detail, how most people in the business world really discounted the importance of search, thinking it was a nice add-on to other core products and services. It was the Google guys that really say the truth here and stuck it out. What things are libraries missing out on due to shortsightedness?

Another thing that really struck me in chapter 8 was the internal struggle as Google experienced explosive growth, to keep the edge and innovative spirit while somehow learning to run the company professionally and keep an eye on organizational issues. It was this 2000-2001 time frame that I mention above.

The next thing to really strike me was the AskJeeves story in chapter 11. Google licensed its ad relevancy software to AskJeeves, to the immense benefit of both companies. I thought this was interesting because two companies that you would think were rivals competing for the same search eyeballs somehow found a way to collaborate and make each of their slices of the pie bigger -- including making a bigger pie. A lesson here for us all -- cooperate and grow or compete and die? Who do we think are our enemies that should be our friends? Google?

Chapter 12 was a big one for me -- where the authors talk about Google's 20% rule. Every employee gets to work one day a week on blue sky projects, things outside the box, the stuff we see in Google Labs. Ultimately, people with ideas have to find others to work on them and to make a case for using more resources than just the 20% time to get the product out the door. But still, the culture of innovation this kind of idea fosters is amazing. Lessons? You bet. Top to bottom, if we want to succeed everyone has to think about innovation and get heard by administrators. Chapter 18 talks about the idea of having a corporate executive chef to make everyone's meals for them. Just creating a environment that's conducive to innovation, no matter what it takes.

On the other hand, a little misinformation is never a bad thing in a business book, especially one on a company with such overpowering ideals. On page 134 talking about the Google News service, the authors quote an engineer that mentions that before Google News, journalists had no way of searching other news sources for information. Of course, we librarians know this is hogwash. Lexis Nexis and its kind aren't free like Google but any journalist working for even a decent sized paper would have had access to it. Sometimes Google would like you to think that only it can provide good information, but sometimes they "ignore" inconvenient truths. Other imperfections that do get some coverage include privacy concerns with advertising in Gmail, censorship of information flowing into China and Google's role in that, some bumps in the road when they went public perhaps betraying an unhealthy arrogance on the part of Brin and Page, the cutthroat nature of the battle with Microsoft, another whiff of arrogance when they talk about Google's role in getting genomic data freely available.

Overall, though, I have to say that this is quite a good book, written in a breezy, journalistic style. Google's story is intimately connected to the story of the early part of the 21st century and we ignore its lessons at our peril: everything is driven by a crazy, intense level of nonstop innovation; search is king; connections between data points can be as important as the data itself, if not more important; share the wealth. Google is a reality, we have to deal with it's implications on our work and personal lives. Its impact is vastly for the better, but that doesn't mean we shouldn't keep an eye on them. Their pride and arrogance can lead to a fall -- putting all our eggs in one basket could be risky.

February 21, 2007

So, you think your institution is change-resistant?

There are institutions where legacy systems can be literally measured in millennia! Imagine being in charge of the Vatican's web sites...and take a look at a video interview with Sister Judith Zoebelein who has that very job. A terrific interview, it really gives a sense of what it's like to bring such an ancient institution into the modern era -- actually, not as hard as it might sound, it seems. I'm really happy I discovered Robert Schoble's show recently; there's lots of cool stuff there.

And while we're on interesting stories with a religious angle, I suggest you read The Story of Sergey Brin: How the Moscow-born entrepreneur cofounded and changed the way the world searches by Mark Malseed. Brin being, of course, co-founder of Google with Larry Page. The story is very interesting in the way it interweaves Brin's Russian-Jewish heritage, his experiences as an immigrant in the US and his drive to build Google from the ground-up. (Via SearchEngineLand.) And speaking of Larry Page, Retrospectacle neatly demolishes his arrogance in trying to tell scientists to be more like toothpaste salespeople. Both these guys have chutzpah to spare, which makes sense when you think about it.

And finally, I'd also like to mention that Frances E. Allen Wins ACM’s Turing Award, the first woman to do so. The Turing Award is the highest honour in the computing field.

From the press release:

ACM, the Association for Computing Machinery, has named Frances E. Allen the recipient of the 2006 A.M. Turing Award for contributions that fundamentally improved the performance of computer programs in solving problems, and accelerated the use of high performance computing. This award marks the first time that a woman has received this honor. The Turing Award, first presented in 1966, and named for British mathematician Alan M. Turing, is widely considered the "Nobel Prize in Computing." It carries a $100,000 prize, with financial support provided by Intel Corporation.

Allen, an IBM Fellow Emerita at the T.J. Watson Research Center, made fundamental contributions to the theory and practice of program optimization, which translates the users' problem-solving language statements into more efficient sequences of computer instructions. Her contributions also greatly extended earlier work in automatic program parallelization, which enables programs to use multiple processors simultaneously in order to obtain faster results. These techniques have made it possible to achieve high performance from computers while programming them in languages suitable to applications. They have contributed to advances in the use of high performance computers for solving problems such as weather forecasting, DNA matching, and national security functions.

February 18, 2007

More videos -- CERN, Google & anti-ID

I just love these science-y videos:

January 17, 2007

Google Librarian Central Blog

The librarians at Google (sort of liaison librarians to librarians) have started their own blog, Google Librarian Central. It seems like a good idea, a way for a bit more two-way communication between the elephant and the mice. I hope that they turn on comments and that they use the blog to really facilitate a dialogue rather than one-way diffusion of information (like the Official Google Blog), reflecting that they can learn from us just as we can take advantage of the great work that they do. It should be very interesting to see how this evolves over the next little while.

Announcing the Librarian Central Blog

I'm pleased to say that today, we're implementing one of your biggest requests. When we asked how we could improve the Google Librarian Newsletter, many of you said, "Make it a blog!" or "Send more up-to-date information." We've taken your feedback to heart, and we're doing just that. Starting today, the Librarian Center will make its home at http://librariancentral.blogspot.com, where you'll find the latest Google news, updates, and tips relevant to the librarian community. The blog includes links to the Newsletter Archive, the Your Stories page, and the Tools and Videos sections. And of course, we'll continue to add to these pages and develop new features.

We're excited about communicating Google's product and feature launches to you as they happen. You can even sign up to receive these blog posts by email, or choose to read them from your Google Personalized Homepage or Google Reader (or your preferred blog reader). For those of you who still prefer to hear from us on a quarterly basis, we'll continue to send out the Librarian Newsletter, which will include the "best of" the previous months' blog posts. As with many Google launches, consider this blog a beta test, open to refinements and changes over time. We'll be looking closely at your feedback, so please let us know what you think.


Update 2007.01.23:
Good news -- and a good sign. They've turned on commenting for the blog.