A Bash DIY data extraction tool
Putting It All Together
You now have the data you need to do your desired analysis. To save typing each command individually, you can put the above commands into a single Bash script as shown in Listing 5.
Listing 5
Complete Bash Script
01 #!/bin/bash 02 # download the websites specified in addresses.txt one by one 03 wget -cv --progress=bar --connect-timeout=30 --force-directories --ignore-length -r -l 7 --convert-links --waitretry=61 -R gif,jpg,png,svg,pdf $(<addresses.txt) 04 # recursively look for the word "abandon" and its variations and print in verbose mode the line before and after the keyword so we can take a quick look at the context 05 grep -r -A1 -B1 "abandon" * > results.txt 06 # find every line that starts with the "--" delimiter and replace it with "12345678" using your favorite text editor 07 # list the first line after "12345678" 08 grep -A 1 -F 12345678 results2.txt > 1stline.txt 09 # delete everything after the "<" character 10 sed 's/<.*//' 1stline.txt > 1stlinefiltered.txt 11 # list every line only once, without its duplicates 12 sort 1stlinefiltered.txt | uniq -u > address_filtered.txt 13 # remove last character form each line (.html-) 14 sed 's/.$//' address_filtered.txt > list_final_address.txt 15 # create a CSV file containing the web addresses 16 cat list_final_address.txt > address.csv 17 # replace "12345678" with "--" in address.csv because "--" might appear in the URL
Then add the addresses to addresses.txt
, each on one line, and save the file in the same folder as the Bash script in Listing 5. Make the script executable with
chmod +x scriptname.sh
Then launch it with ./scriptname.sh
.
With a few simple Bash commands, you have a DIY text data collection tool that delivers a CSV file for use in your favorite statistical application.
« Previous 1 2 3
Buy this article as PDF
(incl. VAT)
Buy Linux Magazine
Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters
Support Our Work
Linux Magazine content is made possible with support from readers like you. Please consider contributing when you've found an article to be beneficial.
News
-
elementary OS 7.1 Now Available for Download
The team behind elementary OS has released the latest version of its operating system with a focus on personalization, inclusivity, accessibility, and privacy.
-
The GNU Project Celebrates Its 40th Birthday
September 27 marks the 40th anniversary of the GNU Project, and it was celebrated with a hacker meeting in Biel/Bienne, Switzerland.
-
Linux Kernel Reducing Long-Term Support
LTS support for the Linux kernel is about to undergo some serious changes that will have a considerable impact on the future.
-
Fedora 39 Beta Now Available for Testing
For fans and users of Fedora Linux, the first beta of release 39 is now available, which is a minor upgrade but does include GNOME 45.
-
Fedora Linux 40 to Drop X11 for KDE Plasma
When Fedora 40 arrives in 2024, there will be a few big changes coming, especially for the KDE Plasma option.
-
Real-Time Ubuntu Available in AWS Marketplace
Anyone looking for a Linux distribution for real-time processing could do a whole lot worse than Real-Time Ubuntu.
-
KSMBD Finally Reaches a Stable State
For those who've been looking forward to the first release of KSMBD, after two years it's no longer considered experimental.
-
Nitrux 3.0.0 Has Been Released
The latest version of Nitrux brings plenty of innovation and fresh apps to the table.
-
Linux From Scratch 12.0 Now Available
If you're looking to roll your own Linux distribution, the latest version of Linux From Scratch is now available with plenty of updates.
-
Linux Kernel 6.5 Has Been Released
The newest Linux kernel, version 6.5, now includes initial support for two very exciting features.