Fuzz introspector
For issues and ideas: https://github.com/ossf/fuzz-introspector/issues

Project functions overview

The following table shows data about each function in the project. The functions included in this table correspond to all functions that exist in the executables of the fuzzers. As such, there may be functions that are from third-party libraries.

For further technical details on the meaning of columns in the below table, please see the Glossary .

Func name Functions filename Args Function call depth Reached by Fuzzers Runtime reached by Fuzzers Combined reached by Fuzzers Fuzzers runtime hit Func lines hit % I Count BB Count Cyclomatic complexity Functions reached Reached by functions Accumulated cyclomatic complexity Undiscovered complexity

Fuzzer details

Fuzzer: extract_text_fuzzer

Call tree

The calltree shows the control flow of the fuzzer. This is overlaid with coverage information to display how much of the potential code a fuzzer can reach is in fact covered at runtime. In the following there is a link to a detailed calltree visualisation as well as a bitmap showing a high-level view of the calltree. For further information about these topics please see the glossary for full calltree and calltree overview

Call tree overview bitmap:

The distribution of callsites in terms of coloring is
Color Runtime hitcount Callsite count Percentage
red 0 464 56.6%
gold [1:9] 0 0.0%
yellow [10:29] 0 0.0%
greenyellow [30:49] 0 0.0%
lawngreen 50+ 355 43.3%
All colors 819 100

Fuzz blockers

The following nodes represent call sites where fuzz blockers occur.

Amount of callsites blocked Calltree index Parent function Callsite Largest blocked function
217 499 pdfminer.pdffont.PDFFont._parse_bbox call site: 00499 pdfminer.pdffont.PDFCIDFont.__init__
93 110 pdfminer.pdftypes.decompress_corrupted call site: 00110 pdfminer.ccitt.ccittfaxdecode
35 719 pdfminer.utils.choplist call site: 00719 pdfminer.pdfinterp.PDFResourceManager.get_font
19 423 pdfminer.encodingdb.name2unicode call site: 00423 .map
17 755 pdfminer.pdfinterp.PDFResourceManager.get_font call site: 00755 pdfminer.pdfinterp.PDFPageInterpreter.init_resources.get_colorspace
13 206 pdfminer.pdftypes.int_value call site: 00206 pdfminer.utils.apply_tiff_predictor
12 444 pdfminer.pdffont.PDFSimpleFont.__init__ call site: 00444 pdfminer.pdftypes.PDFStream.get_data
5 284 pdfminer.psparser.literal_name call site: 00284 pdfminer.pdftypes.int_value
5 408 pdfminer.pdffont.PDFSimpleFont.__init__ call site: 00408 pdfminer.encodingdb.EncodingDB.get_encoding
5 414 pdfminer.encodingdb.EncodingDB.get_encoding call site: 00414 pdfminer.encodingdb.name2unicode
4 235 pdfminer.utils.apply_png_predictor call site: 00235 .enumerate
4 354 pdfminer.pdfinterp.PDFPageInterpreter.process_page call site: 00354 pdfminer.pdfdevice.TagExtractor._write

Runtime coverage analysis

Covered functions
372
Functions that are reachable but not covered
190
Reachable functions
289
Percentage of reachable functions covered
34.26%
NB: The sum of covered functions and functions that are reachable but not covered need not be equal to Reachable functions . This is because the reachability analysis is an approximation and thus at runtime some functions may be covered that are not included in the reachability analysis. This is a limitation of our static analysis capabilities.
Warning: The number of covered functions are larger than the number of reachable functions. This means that there are more functions covered at runtime than are extracted using static analysis. This is likely a result of the static analysis component failing to extract the right call graph or the coverage runtime being compiled with sanitizers in code that the static analysis has not analysed. This can happen if lto/gold is not used in all places that coverage instrumentation is used.
Function name source code lines source lines hit percentage hit

Files reached

filename functions hit
/ 1
...pdfminer.six.fuzzing.extract_text_fuzzer 7
pdfminer.high_level 10
pdfminer.layout 8
pdfminer.utils 21
pdfminer.converter 10
pdfminer.pdfinterp 43
pdfminer.pdfpage 29
pdfminer.pdfparser 3
pdfminer.psparser 7
pdfminer.pdfdocument 48
pdfminer.pdftypes 31
pdfminer.lzw 12
pdfminer.ascii85 7
pdfminer.runlength 5
pdfminer.ccitt 32
pdfminer.pdfdevice 4
pdfminer.pdffont 74
pdfminer.encodingdb 17
pdfminer.cmapdb 29
pdfminer.casting 7

Fuzzer: extract_text_to_fp_fuzzer

Call tree

The calltree shows the control flow of the fuzzer. This is overlaid with coverage information to display how much of the potential code a fuzzer can reach is in fact covered at runtime. In the following there is a link to a detailed calltree visualisation as well as a bitmap showing a high-level view of the calltree. For further information about these topics please see the glossary for full calltree and calltree overview

Call tree overview bitmap:

The distribution of callsites in terms of coloring is
Color Runtime hitcount Callsite count Percentage
red 0 532 60.5%
gold [1:9] 0 0.0%
yellow [10:29] 0 0.0%
greenyellow [30:49] 0 0.0%
lawngreen 50+ 346 39.4%
All colors 878 100

Fuzz blockers

The following nodes represent call sites where fuzz blockers occur.

Amount of callsites blocked Calltree index Parent function Callsite Largest blocked function
217 541 pdfminer.pdffont.PDFFont._parse_bbox call site: 00541 pdfminer.pdffont.PDFCIDFont.__init__
93 152 pdfminer.pdftypes.decompress_corrupted call site: 00152 pdfminer.ccitt.ccittfaxdecode
41 21 pdfminer.converter.PDFConverter._is_binary_stream call site: 00021 pdfminer.converter.HTMLConverter.__init__
35 761 pdfminer.utils.choplist call site: 00761 pdfminer.pdfinterp.PDFResourceManager.get_font
21 856 pdfminer.converter.PDFLayoutAnalyzer.end_page call site: 00856 pdfminer.converter.HTMLConverter.close
19 465 pdfminer.encodingdb.name2unicode call site: 00465 .map
17 797 pdfminer.pdfinterp.PDFResourceManager.get_font call site: 00797 pdfminer.pdfinterp.PDFPageInterpreter.init_resources.get_colorspace
13 248 pdfminer.pdftypes.int_value call site: 00248 pdfminer.utils.apply_tiff_predictor
12 486 pdfminer.pdffont.PDFSimpleFont.__init__ call site: 00486 pdfminer.pdftypes.PDFStream.get_data
11 0 EP call site: 00000 pdfminer.high_level.extract_text_to_fp
5 327 pdfminer.psparser.literal_name call site: 00327 pdfminer.pdftypes.int_value
5 450 pdfminer.pdffont.PDFSimpleFont.__init__ call site: 00450 pdfminer.encodingdb.EncodingDB.get_encoding

Runtime coverage analysis

Covered functions
372
Functions that are reachable but not covered
213
Reachable functions
312
Percentage of reachable functions covered
31.73%
NB: The sum of covered functions and functions that are reachable but not covered need not be equal to Reachable functions . This is because the reachability analysis is an approximation and thus at runtime some functions may be covered that are not included in the reachability analysis. This is a limitation of our static analysis capabilities.
Warning: The number of covered functions are larger than the number of reachable functions. This means that there are more functions covered at runtime than are extracted using static analysis. This is likely a result of the static analysis component failing to extract the right call graph or the coverage runtime being compiled with sanitizers in code that the static analysis has not analysed. This can happen if lto/gold is not used in all places that coverage instrumentation is used.
Function name source code lines source lines hit percentage hit

Files reached

filename functions hit
/ 1
...pdfminer.six.fuzzing.extract_text_to_fp_fuzzer 11
pdfminer.high_level 16
pdfminer.image 2
pdfminer.converter 31
pdfminer.pdfdevice 5
pdfminer.pdfinterp 43
pdfminer.pdfpage 29
pdfminer.pdfparser 3
pdfminer.psparser 7
pdfminer.pdfdocument 47
pdfminer.pdftypes 31
pdfminer.lzw 12
pdfminer.ascii85 7
pdfminer.runlength 5
pdfminer.ccitt 32
pdfminer.utils 17
pdfminer.layout 6
pdfminer.pdffont 74
pdfminer.encodingdb 17
pdfminer.cmapdb 29
pdfminer.casting 7

Fuzzer: page_extraction_fuzzer

Call tree

The calltree shows the control flow of the fuzzer. This is overlaid with coverage information to display how much of the potential code a fuzzer can reach is in fact covered at runtime. In the following there is a link to a detailed calltree visualisation as well as a bitmap showing a high-level view of the calltree. For further information about these topics please see the glossary for full calltree and calltree overview

Call tree overview bitmap:

The distribution of callsites in terms of coloring is
Color Runtime hitcount Callsite count Percentage
red 0 490 52.2%
gold [1:9] 0 0.0%
yellow [10:29] 0 0.0%
greenyellow [30:49] 0 0.0%
lawngreen 50+ 448 47.7%
All colors 938 100

Fuzz blockers

The following nodes represent call sites where fuzz blockers occur.

Amount of callsites blocked Calltree index Parent function Callsite Largest blocked function
217 488 pdfminer.pdffont.PDFFont._parse_bbox call site: 00488 pdfminer.pdffont.PDFCIDFont.__init__
93 111 pdfminer.pdftypes.decompress_corrupted call site: 00111 pdfminer.ccitt.ccittfaxdecode
34 708 pdfminer.utils.choplist call site: 00708 pdfminer.pdfinterp.PDFResourceManager.get_font
19 412 pdfminer.encodingdb.name2unicode call site: 00412 .map
16 743 pdfminer.pdfinterp.PDFResourceManager.get_font call site: 00743 pdfminer.pdfinterp.PDFPageInterpreter.init_resources.get_colorspace
13 207 pdfminer.pdftypes.int_value call site: 00207 pdfminer.utils.apply_tiff_predictor
12 433 pdfminer.pdffont.PDFSimpleFont.__init__ call site: 00433 pdfminer.pdftypes.PDFStream.get_data
10 804 pdfminer.layout.LTLayoutContainer.analyze call site: 00804 obj0.is_voverlap
6 0 EP call site: 00000 pdfminer.high_level.extract_pages
5 286 pdfminer.psparser.literal_name call site: 00286 pdfminer.pdftypes.int_value
5 397 pdfminer.pdffont.PDFSimpleFont.__init__ call site: 00397 pdfminer.encodingdb.EncodingDB.get_encoding
5 403 pdfminer.encodingdb.EncodingDB.get_encoding call site: 00403 pdfminer.encodingdb.name2unicode

Runtime coverage analysis

Covered functions
372
Functions that are reachable but not covered
226
Reachable functions
348
Percentage of reachable functions covered
35.06%
NB: The sum of covered functions and functions that are reachable but not covered need not be equal to Reachable functions . This is because the reachability analysis is an approximation and thus at runtime some functions may be covered that are not included in the reachability analysis. This is a limitation of our static analysis capabilities.
Warning: The number of covered functions are larger than the number of reachable functions. This means that there are more functions covered at runtime than are extracted using static analysis. This is likely a result of the static analysis component failing to extract the right call graph or the coverage runtime being compiled with sanitizers in code that the static analysis has not analysed. This can happen if lto/gold is not used in all places that coverage instrumentation is used.
Function name source code lines source lines hit percentage hit

Files reached

filename functions hit
/ 1
...pdfminer.six.fuzzing.page_extraction_fuzzer 8
pdfminer.high_level 9
pdfminer.layout 68
pdfminer.utils 27
pdfminer.converter 13
pdfminer.pdfinterp 41
pdfminer.pdfpage 27
pdfminer.pdfparser 3
pdfminer.psparser 7
pdfminer.pdfdocument 47
pdfminer.pdftypes 32
pdfminer.lzw 12
pdfminer.ascii85 7
pdfminer.runlength 5
pdfminer.ccitt 32
pdfminer.pdfdevice 4
pdfminer.pdffont 75
pdfminer.encodingdb 17
pdfminer.cmapdb 29
pdfminer.casting 7

Analyses and suggestions

Optimal target analysis

Remaining optimal interesting functions

The following table shows a list of functions that are optimal targets. Optimal targets are identified by finding the functions that in combination, yield a high code coverage.

Func name Functions filename Arg count Args Function depth hitcount instr count bb count cyclomatic complexity Reachable functions Incoming references total cyclomatic complexity Unreached complexity
pdfminer.converter.HTMLConverter.receive_layout.render pdfminer.converter 1 ['N/A'] 6 0 24 15 9 75 2 244 196
pdfminer.psparser.PSStackParser.nextobject pdfminer.psparser 1 ['N/A'] 3 0 15 13 8 48 0 167 127
pdfminer.pdfdocument.PDFStandardSecurityHandlerV5.authenticate pdfminer.pdfdocument 2 ['N/A', 'N/A'] 4 0 0 2 4 26 0 85 82
pdfminer.pdfinterp.PDFPageInterpreter.do_TJ pdfminer.pdfinterp 2 ['N/A', 'N/A'] 3 0 2 2 4 31 3 103 64
pdfminer.pdfinterp.PDFPageInterpreter.do_Do pdfminer.pdfinterp 2 ['N/A', 'N/A'] 6 0 8 3 4 41 0 131 52
pdfminer.cmapdb.CMapParser.do_keyword pdfminer.cmapdb 3 ['N/A', 'N/A', 'N/A'] 4 0 27 29 15 50 0 176 39
pdfminer.pdfdocument.PageLabels.labels pdfminer.pdfdocument 1 ['N/A'] 3 0 2 3 4 24 0 82 38
pdfminer.pdfdocument.PDFDocument.getobj pdfminer.pdfdocument 2 ['N/A', 'N/A'] 7 0 4 6 5 105 2 361 37

Implementing fuzzers that target the above functions will improve reachability such that it becomes:

Functions statically reachable by fuzzers
47.0%
289 / 609
Cyclomatic complexity statically reachable by fuzzers
50.0%
1057 / 2105

All functions overview

If you implement fuzzers for these functions, the status of all functions in the project will be:

Func name Functions filename Args Function call depth Reached by Fuzzers Runtime reached by Fuzzers Combined reached by Fuzzers Fuzzers runtime hit Func lines hit % I Count BB Count Cyclomatic complexity Functions reached Reached by functions Accumulated cyclomatic complexity Undiscovered complexity

Fuzz engine guidance

This sections provides heuristics that can be used as input to a fuzz engine when running a given fuzz target. The current focus is on providing input that is usable by libFuzzer.

fuzzing/extract_text_fuzzer.py

Dictionary

Use this with the libFuzzer -dict=DICT.file flag


Fuzzer function priority

Use one of these functions as input to libfuzzer with flag: -focus_function name

-focus_function=['pdfminer.pdffont.PDFFont._parse_bbox', 'pdfminer.pdftypes.decompress_corrupted', 'pdfminer.utils.choplist', 'pdfminer.encodingdb.name2unicode', 'pdfminer.pdfinterp.PDFResourceManager.get_font', 'pdfminer.pdftypes.int_value', 'pdfminer.pdffont.PDFSimpleFont.__init__', 'pdfminer.psparser.literal_name', 'pdfminer.encodingdb.EncodingDB.get_encoding']

fuzzing/extract_text_to_fp_fuzzer.py

Dictionary

Use this with the libFuzzer -dict=DICT.file flag


Fuzzer function priority

Use one of these functions as input to libfuzzer with flag: -focus_function name

-focus_function=['pdfminer.pdffont.PDFFont._parse_bbox', 'pdfminer.pdftypes.decompress_corrupted', 'pdfminer.converter.PDFConverter._is_binary_stream', 'pdfminer.utils.choplist', 'pdfminer.converter.PDFLayoutAnalyzer.end_page', 'pdfminer.encodingdb.name2unicode', 'pdfminer.pdfinterp.PDFResourceManager.get_font', 'pdfminer.pdftypes.int_value', 'pdfminer.pdffont.PDFSimpleFont.__init__']

fuzzing/page_extraction_fuzzer.py

Dictionary

Use this with the libFuzzer -dict=DICT.file flag


Fuzzer function priority

Use one of these functions as input to libfuzzer with flag: -focus_function name

-focus_function=['pdfminer.pdffont.PDFFont._parse_bbox', 'pdfminer.pdftypes.decompress_corrupted', 'pdfminer.utils.choplist', 'pdfminer.encodingdb.name2unicode', 'pdfminer.pdfinterp.PDFResourceManager.get_font', 'pdfminer.pdftypes.int_value', 'pdfminer.pdffont.PDFSimpleFont.__init__', 'pdfminer.layout.LTLayoutContainer.analyze', 'pdfminer.psparser.literal_name']

Files and Directories in report

This section shows which files and directories are considered in this report. The main reason for showing this is fuzz introspector may include more code in the reasoning than is desired. This section helps identify if too many files/directories are included, e.g. third party code, which may be irrelevant for the threat model. In the event too much is included, fuzz introspector supports a configuration file that can exclude data from the report. See the following link for more information on how to create a config file: link

Files in report

Source file Reached by Covered by
pdfminer.high_level ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
[] []
pdfminer.converter ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
logging [] []
contextlib [] []
unittest [] []
json [] []
re [] []
atheris [] []
itertools [] []
[] []
pdfminer.fontmetrics [] []
pdfminer.psparser ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.layout ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.utils ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
io [] []
pdfminer.pdffont ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
typing [] []
os [] []
pdfminer.encodingdb ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
...pdfminer.six.fuzzing.extract_text_to_fp_fuzzer ['extract_text_to_fp_fuzzer'] []
pygame [] []
pdfminer.pdfcolor [] []
hashlib [] []
pdfminer.casting ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.pdfinterp ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.pdfpage ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.pdftypes ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
unicodedata [] []
pdfminer.ccitt ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.pdfdevice ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.jbig2 [] []
importlib [] []
pdfminer.cmapdb ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.arcfour [] []
pdfminer.glyphlist [] []
PIL [] []
math [] []
pdfminer [] []
struct [] []
gzip [] []
pdfminer.psexceptions [] []
array [] []
heapq [] []
html [] []
fuzzing [] []
...pdfminer.six.fuzzing.extract_text_fuzzer ['extract_text_fuzzer'] []
pdfminer.pdfexceptions [] []
base64 [] []
zlib [] []
cryptography [] []
pdfminer.ascii85 ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer._saslprep [] []
stringprep [] []
pdfminer.data_structures [] []
charset_normalizer [] []
warnings [] []
pdfminer.pdfdocument ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
binascii [] []
pdfminer.settings [] []
pdfminer.runlength ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
collections [] []
pdfminer.lzw ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.pdfparser ['extract_text_fuzzer', 'extract_text_to_fp_fuzzer', 'page_extraction_fuzzer'] []
pdfminer.image ['extract_text_to_fp_fuzzer'] []
pdfminer.latin_enc [] []
...pdfminer.six.fuzzing.page_extraction_fuzzer ['page_extraction_fuzzer'] []

Directories in report

Directory