CHF 135.00

Ulrich Matter

An Introduction to Web Mining
with Applications in R

English · Paperback / Softback

Shipping usually within 6 to 7 weeks

Description

This book is devoted to the art and science of web mining showing how the world's largest information source can be turned into structured, research-ready data. Drawing on many years of teaching graduate courses on Web Mining and on numerous large-scale research projects in web mining contexts, the author provides clear explanations of key web technologies combined with hands-on R tutorials that work in the real world and keep working as the web evolves.
Through the book, readers will learn how to
- scrape static and dynamic/JavaScript-heavy websites
- use web APIs for structured data extraction from web sources
- build fault-tolerant crawlers and cloud-based scraping pipelines
- navigate CAPTCHAs, rate limits, and authentication hurdles
- integrate AI-driven tools to speed up every stage of the workflow
- apply ethical, legal, and scientific guidelines to their web mining activities
Part I explains why web data matters and leads the reader through a first hello-scrape in R while introducing HTML, HTTP, and CSS. Part II explores how the modern web works and shows, step by step, how to move from scraping static pages to collecting data from APIs and JavaScript-driven sites. Part III focuses on scaling up: building reliable crawlers, dealing with log-ins and CAPTCHAs, using cloud resources, and adding AI helpers. Part IV looks at ethical, legal, and research standards, offering checklists and case studies, enabling the reader to make responsible choices. Together, these parts give a clear path from small experiments to large-scale projects.
This valuable guide is written for a wide readership from graduate students taking their first steps in data science to seasoned researchers and analysts in economics, social science, business, and public policy. It will be a lasting reference for anyone with an interest in extracting insight from the web whether working in academia, industry, or the public sector.

About the author

Ulrich Matter is Professor of Applied Data Science at Bern University of Applied Sciences and Affiliate Professor of Economics at the University of St. Gallen. His primary research interests lie at the intersection of data science, political economics, and media economics.

Summary

This book is devoted to the art and science of web mining — showing how the world's largest information source can be turned into structured, research-ready data. Drawing on many years of teaching graduate courses on Web Mining and on numerous large-scale research projects in web mining contexts, the author provides clear explanations of key web technologies combined with hands-on R tutorials that work in the real world — and keep working as the web evolves.
Through the book, readers will learn how to

- scrape static and dynamic/JavaScript-heavy websites

- use web APIs for structured data extraction from web sources

- build fault-tolerant crawlers and cloud-based scraping pipelines

- navigate CAPTCHAs, rate limits, and authentication hurdles

- integrate AI-driven tools to speed up every stage of the workflow

- apply ethical, legal, and scientific guidelines to their web mining activities

Part I explains why web data matters and leads the reader through a first “hello-scrape” in R while introducing HTML, HTTP, and CSS. Part II explores how the modern web works and shows, step by step, how to move from scraping static pages to collecting data from APIs and JavaScript-driven sites. Part III focuses on scaling up: building reliable crawlers, dealing with log-ins and CAPTCHAs, using cloud resources, and adding AI helpers. Part IV looks at ethical, legal, and research standards, offering checklists and case studies, enabling the reader to make responsible choices. Together, these parts give a clear path from small experiments to large-scale projects.
This valuable guide is written for a wide readership — from graduate students taking their first steps in data science to seasoned researchers and analysts in economics, social science, business, and public policy. It will be a lasting reference for anyone with an interest in extracting insight from the web — whether working in academia, industry, or the public sector.

Product details

Authors	Ulrich Matter
Publisher	Springer, Berlin

Content	Book
Product form	Paperback / Softback
Publication date	26.08.2025
Subject	Natural sciences, medicine, IT, technology > Mathematics > Probability theory, stochastic theory, mathematica

EAN	9783031966378
ISBN	978-3-0-3196637-8
Pages	251
Illustrations	XXI, 251 p. 34 illus., 18 illus. in color.
Dimensions (packing)	15.5 x 1.5 x 23.5 cm
Weight (packing)	423 g

Series	Use R!
Subjects	Data Mining, Web Scraping, Sozialforschung und -statistik, Wissensbasierte Systeme, Expertensysteme, Restful apis, Data Mining and Knowledge Discovery, Methodology of Data Collection and Processing, r programming, Web Mining, tidyverse, JSON Processing, Cloud Scraping, AJAX Handling, Headless Browsers, XML Parsing, CSS Selectors, JavaScript Scraping, AI-assisted scraping, Pagination Handling, API Data Extraction, Dynamic Websites, Web Crawler Design, HTML Parsing, reproducible research, HTTP Requests, Wirtschaftswissenschaft, Finanzen, Betriebswirtschaft und Management

Customer reviews

No reviews have been written for this item yet. Write the first review and be helpful to other users when they decide on a purchase.

Write a review

Thumbs up or thumbs down? Write your own review.

Your contact at CeDe