Vi-DIFF: Understanding Web Pages Changes

9 years 9 months ago
Vi-DIFF: Understanding Web Pages Changes
Nowadays, many applications are interested in detecting and discovering changes on the web to help users to understand page updates and more generally, the web dynamics. Web archiving is one of these fields where detecting changes on web pages is important. Archiving institutes are collecting and preserving different web site versions for future generation. A major problem encountered by archiving systems is to understand what happened between two versions of web pages. In this paper, we address this requirement by proposing a new change detection approach that computes the semantic differences between two versions of HTML web pages. Our approach, called Vi-DIFF, detects changes on the visual representation of web pages. It detects two types of changes: content and structural changes. Content changes include modifications on text, hyperlinks and images. In contrast, structural changes alter the visual appearance of the page and the structure of its blocks. Our ViDIFF solution can s...
Zeynep Pehlivan, Myriam Ben Saad, Stéphane
Added 24 Jan 2011
Updated 24 Jan 2011
Type Journal
Year 2010
Where DEXA
Authors Zeynep Pehlivan, Myriam Ben Saad, Stéphane Gançarski
Comments (0)