Distributed Lucene for Hadoop

19th August 2008 in London at Sekforde Street

This talk describes a parallel, distributed free text index written at HP Labs Bristol called Distributed Lucene. Distributed Lucene is based on two Apache open source projects, Hadoop and Lucene, and follows a design originally proposed by Doug Cutting. It was written to gain a better understanding of the Apache Hadoop architecture, and to investigate approaches to creating large, scalable free text indexes. For more information see the accompanying HP Labs technical report.

Mark Butler

