html2lcov - Recover lcov data from a genhtml HTML report.
- Manual section:
1
- Manual group:
LCOV Tools
NAME
- html2lcov
Recover lcov
.infocoverage data and a source diff from an HTML report
SYNOPSIS
html2lcov [--output filename] [--source-directory dir]+ [--current-file file]+ [options] report_directory+
DESCRIPTION
html2lcov screen-scrapes a genhtml-generated HTML coverage report to
recover:
The LCOV
.infocoverage data that was used to produce the report.genhtmlcopies its input.infofiles into the top level of the report directory when it is run with--save;html2lcovfinds and aggregates those files, using the same layout definition thatgenhtml --saveuses to write them (so the reader and writer cannot drift). It is an error if the report does not contain any such.infofile.The source code that was embedded in the per-file HTML pages. This is compared against the source found under
--source-directoryto produce a universal diff and a human-readable difference report. Note thathtml2lcovneeds to see all of the original source code: the input HTML report must not be a subset report containing only part of the source code (i.e., not a report generated using a--select-scriptcallback).
The source directory may contain new files that were not in the report, and may omit files that were in the report; these differences are documented in the difference report rather than being treated as errors. It is an error if the source directory contains no file that matches any file in the report.
Because genhtml expands tabs and normalizes some whitespace when it renders
source into HTML, the comparison is run with diff -b, which ignores
differences in the amount of whitespace (a tab and any run of spaces compare
equal), so re-indented or tab-vs-space-only changes do not appear as spurious
differences.
USE CASES
The common use case for html2lcov is when the user has code and/or
test changes in their sandbox, has a coverage report from earlier in the day,
and now wants to see whether subsequent changes are adequately tested.
Unfortunately: the user isn't entirely sure what the code looked like
when the report was generated - and does not have a corresponding label.
In that case: file diffs and baseline coverage data
can be extracted from the report, and then used as the genhtml
--diff-file and --baseline-file inputs.
Another common case is when some or all of the source code is generated and not under revision control - but the user wants to verify that it is properly tested after configuration and/or generator changes.
OPTIONS
html2lcov supports the options that are common to the other tools in the
LCOV suite, and honours the corresponding lcovrc(5) settings. In
particular:
file selection and path manipulation:
--include,--exclude,--erase-functions,--substitute,--filter,--omit-lines,--demangle-cpp. These apply to the coverage data recovered from the report exactly as they do when the same data is read bylcov- so a file which is excluded is also absent from the universal diff and from the difference report.parallelism and resource limits:
--parallel(-j),--memory,--tempdir,--preserve.html2lcovforks children both to read the report's saved.infofile(s) and to scrape and diff the report's source pages, so--parallelspeeds up both phases of a large report; see PARALLELISM.message control and diagnostics:
--ignore-errors,--keep-going,--expect-message-count,--msg-log,--quiet,--verbose,--debug.configuration and provenance:
--config-file,--rc,--comment,--version,--help.
The tool-specific options are:
--outputfilename,-ofilenameName a base for the output files. Any extension is stripped from filename to form base, and each artifact then appends its own extension: the aggregated coverage is written to base
.info, the human-readable difference report to base.rpt, the universal diff to base.udiff(unless--diff-fileredirects it), and the profile (if--profileis used) to base.json. When--outputis omitted base ishtml2lcovand the universal diff is written to standard output.--source-directorydirSearch dir for the current source files to diff against the report's embedded source. May be used more than once. Files are matched by their path relative to the report's recovered source root. When
--source-directoryis not specified it defaults to.(the current directory).--current-filefilenameAn lcov
.infofile whoseSF:records name the current source files. May be specified more than once, and each value may itself be a list separated by the configured list separator (seelcovrc(5)).When any
--current-fileis given, its union ofSF:paths becomes an authoritative whitelist of the current source files:Only files named in a
SF:record may appear in the universal diff (as changed or added). Every other file found under--source-directoryis ignored.A file present in the report (baseline) but not in the current set is treated as removed.
A file present in the current set but not in the report is treated as added.
html2lcovuses the.infoonly as a file list; its coverage data is not merged into any output. The expected workflow is to feed thehtml2lcovuniversal diff togenhtml --diff-fileand the aggregated.infooutput togenhtml --baseline-file, using this same--current-file.infoas thegenhtmlcurrent input.A current
SF:path is matched to a report file when it equals either the report file's recovered path or its report-root-relative path (after any--substituterules are applied); relative and absolute path styles are preserved on output. If a file appears in both sets under names that differ only in path (a shared basename or path tail),html2lcovemits an ignorablemismatchwarning. Use--substitute(or an external tool such as sed(1)) to reconcile path differences the tool cannot resolve on its own.If a
SF:path names a file that cannot be read from disk under--source-directory,html2lcovreports an ignorablesourceerror; when that error is ignored, the file is dropped from the universal diff and its record is removed from the aggregated.infooutput.A missing
--current-fileis a fatal error; an empty file raises an ignorableemptyerror; a non-empty file that contains noSF:record is not an LCOV file and is a fatal error.--diff-filefilenameWrite the universal diff to filename instead of the default base
.udiff(or standard output when--outputis omitted).--profile[filename]Write timing and statistics data as JSON to filename (default is base
.json).
See lcov(1) and lcovrc(5) for details of other supported options and configuration settings.
OUTPUTS
html2lcov produces up to four distinct artifacts:
universal diff (base
.udiff) - written to base.udiff, or to standard output when--outputis omitted;--diff-fileredirects it elsewhere. Standarddiff -uformat, so it can be consumed bygenhtml --diff. Files only under--source-directoryappear as additions (versus/dev/null); files only in the report appear as removals. May be empty when there are no code changes. For a changed file both diff headers (---and+++) name the same path -- the source path recovered from the report, which is exactly theSF:record in the aggregated.info-- sogenhtmlcan associate the diff with the recovered baseline coverage.aggregated coverage (base
.info) - thecurrentcoverage recovered from the report's saved.infofile(s).difference report (base
.rpt) - a human-readable table listing every changed, added, removed, and unchanged file, with added/deleted line counts.profile (base
.json) - timing/statistics data, when--profileis used.
PARALLELISM
Both of the expensive phases of html2lcov are parallelized, and both are
governed by --parallel (-j) and --memory:
reading the
.infofile(s) saved in the report - the same parallel parse thatlcov --add-tracefileuses.scraping the source out of the report's HTML pages and diffing it against
--source-directory. Each child scrapes and diffs a group of files and returns only the classification and the udiff text, so the recovered source never accumulates in the parent. Files are grouped by the size of the HTML page each is scraped from, so that the groups cost roughly the same; a report with only a handful of files is processed in this process rather than paying for a fork.
The output does not depend on --parallel: the .info, .udiff and
.rpt artifacts are identical for any value.
EXAMPLES
Recover coverage and diff current source against a saved report:
# Generate coverage and an HTML report, saving the input .info files
$ lcov --capture -d . -o cov.info
$ genhtml cov.info --save -o html_report
# make code and/or test changes
...
# capture new coverage data
$ lcov --capture -d . -o current.info
# Recover the .info data and diff the current source tree
$ html2lcov -o changes html_report --source-directory ./src --current-file current.info
# 'changes.udiff' is the universal diff (diff -u format)
# 'changes.info' is the recovered coverage data
# 'changes.rpt' is the human-readable difference report
# now generate a differential report - to see what has changed:
$ genhtml -o differential --baseline-file changes.info --diff-file changes.udiff current.info
Also see the example_html2lcov in the example directory |ToolName|_HOME/share/lcov/example:
$ cd |ToolName|_HOME/share/lcov/example
$ make example_html2lcov
NOTES
Coverage cannot be recovered from a report that was built without --save,
because the .info files are not present in the report directory.
html2lcov treats this as an error rather than attempting a lossy scrape of
the coverage counts from the HTML.
If the source under --source-directory is identical to the source embedded
in the report, there are no differences to emit. This is a normal, expected
outcome (for example, when the current source has not changed since the report
was built), so html2lcov reports it as a non-fatal empty warning
(No source code differences found) and still exits successfully. Use
--ignore-errors empty to suppress the warning entirely.