Class ParsedElement
One region of a parsed page: its role, where it sits, its place in the reading order, and its typed content.
public sealed class ParsedElement
- Inheritance
-
ParsedElement
- Inherited Members
Remarks
The concrete type of Content tells how its text is encoded: TextContent (Markdown), TableContent (HTML), FormulaContent (LaTeX), ChartContent (the data a chart plots, as a Markdown table), FigureContent (a picture), and the text kinds that belong to a figure or repeat another element (FigureCaptionContent, FigureLabelContent, RepeatedTextContent).
Properties
- Category
The region's semantic role.
- Confidence
Confidence in the region and its role, in [0, 1]; 1 when none is estimated.
- Content
The typed content of the region.
- PrintedWords
The words printed inside a figure's box, in print order: the document's own text where it has any, else the words read off the figure's pixels. Empty for every other element and for a figure that prints no word. They describe the figure to a consumer that needs its text (a layout view, a search index, an alternative text) and are never part of its rendering.
- ReadingIndex
Zero-based position in the page's reading order.
Methods
- GetImage()
The picture of this figure or chart as the page shows it, when the parser keeps figure images (IncludeFigureImages).
- ToHtml()
Renders the element as an HTML fragment, as it appears in its page's rendering (ToHtml()): a heading, a paragraph, a table, a formula, a figure or a chart, carrying its category and bounds. Its captions, which the page rendering attaches to it, are elements of their own here. Empty for an element the page rendering leaves out.
- ToJson()
Serializes the element as JSON, in the same schema as one entry of a page's
elementsarray in ToJson().
- ToMarkdown()
Renders the element as Markdown, as it appears in its page's rendering: a heading with its marks, a table as HTML, a formula fenced, a chart as its data table under its caption and legend. Empty for an element the page rendering leaves out.