Indexing and searching text with Lucene
Smart Search
Even state-of-the-art computers need to use clever methods to process ever-increasing amounts of document data. The open source Lucene framework uses inverted indexing for fast searches of document collections.
Nowadays, almost any commercially available hard drive can store more text than a whole library. In the digital world, a traditional system such as a card catalog or a knowledgeable librarian is no longer adequate to help find the right shelf. Even software equivalents such as find or zgrep are not always fast enough to track a particular piece of information amongst giga- or terabytes of data.
The science that deals with this type of search problem is called information retrieval. Computer scientists have developed sophisticated methods for tracking down files that users don’t even know exist. The free Java library Lucene implements some of these methods. Doug Cutting published an early version of Lucene in 1999. Two years later, the project, which carries the middle name of Cutting’s wife, came under the auspices of the Apache Foundation when it joined the Apache Jakarta Project. Lucene has been available in Version 4.0 since October 2012. The index file structures are backward compatible, so the transition from 3.6 to 4.0 does not cause any problems. Over the years, Lucene has become one of the most widely used solutions for indexing and searching text. (See the box titled “Lucene In All Its Facets.”)
Buy this article as PDF
(incl. VAT)
Buy Linux Magazine
Direct Download
Read full article as PDF:
Price $2.95
Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters
News
-
An All-Snap Version of Ubuntu is In The Works
Along with the standard deb version of the open-source operating system, Canonical will release an-all snap version.
-
Mageia 9 Beta 2 Ready for Testing
The latest beta of the popular Mageia distribution now includes the latest kernel and plenty of updated applications.
-
KDE Plasma 6 Looks to Bring Basic HDR Support
The KWin piece of KDE Plasma now has HDR support and color management geared for the 6.0 release.
-
Bodhi Linux 7.0 Beta Ready for Testing
The latest iteration of the Bohdi Linux distribution is now available for those who want to experience what's in store and for testing purposes.
-
Changes Coming to Ubuntu PPA Usage
The way you manage Personal Package Archives will be changing with the release of Ubuntu 23.10.
-
AlmaLinux 9.2 Now Available for Download
AlmaLinux has been released and provides a free alternative to upstream Red Hat Enterprise Linux.
-
An Immutable Version of Fedora Is Under Consideration
For anyone who's a fan of using immutable versions of Linux, the Fedora team is currently considering adding a new spin called Fedora Onyx.
-
New Release of Br OS Includes ChatGPT Integration
Br OS 23.04 is now available and is geared specifically toward web content creation.
-
Command-Line Only Peropesis 2.1 Available Now
The latest iteration of Peropesis has been released with plenty of updates and introduces new software development tools.
-
TUXEDO Computers Announces InfinityBook Pro 14
With the new generation of their popular InfinityBook Pro 14, TUXEDO upgrades its ultra-mobile, powerful business laptop with some impressive specs.