Wednesday, November 18, 2009

Automated Data Processing: Too Big for Our Puny Brains

Over my past few blog entries I've been thinking more and more about automated processing of scientific data, as well as distributed efforts to handle massive backlogs of observations. Science is now generating such huge data sets that our ability to intelligently query them in any manual fashion is impossible. For example, CERN's Large Hadron Collider will be producing 40 TB of data every single day. For these reasons, if we hope to make any progress with scientific data, we'll need to employ AI (artificial intelligence) as our new scientific tools. Maybe this topic will seem to be at the far reaches of what we've been discussing in class, but I think it is a useful look at just how these datasets will actually be used. Instead of human searchers crawling the databases, we'll see automated systems trying to distill the knowledge from data bits.

An article from this month's Communications of the ACM describes the efforts of two computer scientists from Cornell University to build a new machine learning system. Whereas older machine learning systems sought to build predictions based on data, this new system was looking for basic invariant relationships such as the conservation of energy, that are scientific constants. This kind of system could formulate scientific laws as basic as the law of gravitation.

The breakthrough in this case comes from giving computers relatively little starting information and instead allowing them to derive rules and theories as they proceed, testing different ones and weighing their success in describing the situation. Other scientists have expressed interest in this system, because it is also scalable across domains. The ACM article on this mentions that this system automates the final part of the traditional scientific paradigm: from observational data to model formulation, to predictions, to laws, to explanatory theories.

Interestingly, the end of the article hints that the results of this system may be beyond human understanding. Interpretation of the results and their meaning may not be possible in human terms. This raises the question, with all this data, what good is it? How does it further our knowledge? How do we even know if the results are correct, if we can't understand them?

No comments:

Post a Comment

Note: Only a member of this blog may post a comment.