可以树.恢复模式下的 XMLParser 仍然抛出解析错误?



我有一个实用程序方法,它使用创建为etree.XMLParser(recover=True)的解析器解析XML。我想在单元测试中测试失败场景。除了空输入抛出lxml.etree.XMLSyntaxError,我似乎无法破坏解析器。

我的问题是:是否可以为此解析器构造StringIOBytesIO输入,以便解析器抛出解析错误?

下面是一些示例(使用 Python 3.5 和 lxml 4.3.3 测试):

from io import BytesIO
from lxml import etree

def parse(xml):
parser = etree.XMLParser(recover=True)
elem = etree.parse(BytesIO(xml), parser)
print(etree.tostring(elem))

parse(b'<broken<')  # prints b'<broken/>'
parse(b'</lf|jf>')  # prints None
parse('<?xml encoding="ascii"?><foo>æøå</foo>'.encode('utf-8'))  # prints b'<foo/>'
parse(b'')  # Throws lxml.etree.XMLSyntaxError

如果我在您显示的任何错误输入的开头打一个 NULL 字符,而不会引发错误,我确实会收到错误。例如:

parse(b'<broken<')

生产:

Traceback (most recent call last):
File "test.py", line 13, in <module>
parse(b'<broken<')  # prints b'<broken/>'
File "test.py", line 9, in parse
elem = etree.parse(BytesIO(xml), parser)
File "src/lxml/etree.pyx", line 3435, in lxml.etree.parse
File "src/lxml/parser.pxi", line 1857, in lxml.etree._parseDocument
File "src/lxml/parser.pxi", line 1877, in lxml.etree._parseMemoryDocument
File "src/lxml/parser.pxi", line 1765, in lxml.etree._parseDoc
File "src/lxml/parser.pxi", line 1127, in lxml.etree._BaseParser._parseDoc
File "src/lxml/parser.pxi", line 601, in lxml.etree._ParserContext._handleParseResultDoc
File "src/lxml/parser.pxi", line 711, in lxml.etree._handleParseResult
File "src/lxml/parser.pxi", line 640, in lxml.etree._raiseParseError
File "<string>", line 1
lxml.etree.XMLSyntaxError: Document is empty, line 1, column 1

不是因为你正在使用 recover=True?

恢复 - 努力解析损坏的 XML

我更改了恢复=假,我得到:

Traceback (most recent call last):
File "./foo.py", line 11, in <module>
parse(b'<broken<')  # prints b'<broken/>'
File "./foo.py", line 7, in parse
elem = etree.parse(BytesIO(xml), parser)
File "src/lxml/etree.pyx", line 3435, in lxml.etree.parse
File "src/lxml/parser.pxi", line 1857, in lxml.etree._parseDocument
File "src/lxml/parser.pxi", line 1877, in lxml.etree._parseMemoryDocument
File "src/lxml/parser.pxi", line 1765, in lxml.etree._parseDoc
File "src/lxml/parser.pxi", line 1127, in lxml.etree._BaseParser._parseDoc
File "src/lxml/parser.pxi", line 601, in lxml.etree._ParserContext._handleParseResultDoc
File "src/lxml/parser.pxi", line 711, in lxml.etree._handleParseResult
File "src/lxml/parser.pxi", line 640, in lxml.etree._raiseParseError
File "<string>", line 1
lxml.etree.XMLSyntaxError: error parsing attribute name, line 1, column 8

我错过了什么吗?

最新更新