Ruby-遍历nokogiri元素

我有一个类似的html：

...
<table>
<tbody>
    ...
    <tr>
    <th> head </th>
    <td> td1 text<td>
    <td> td2 text<td>
    ...
    </tr>
</tbody>
<tfoot>
</tfoot>
</table>
...

我用的是带红宝石的野宫。我想遍历每一行，并将th的文本和相应的td放入哈希中。

require "nokogiri"
#Parses your HTML input
html_data = "...stripped HTML markup code..." 
html_doc = Nokogiri::HTML html_data
#Iterates over each row in your table
#Note that you may need to clarify the CSS selector below
result = html_doc.css("table tr").inject({}) do |all, row|
  #Modify if you need to collect only the first td, for example
  all[row.css("th").text] = row.css("td").text
end

我没有运行这段代码，所以我不能完全确定，但总体想法应该是正确的：

html_doc = Nokogiri::HTML("<html> ... </html>")
result = []
html_doc.xpath("//tr").each do |tr|
  hash = {}
  tr.children.each do |node|
    hash[node.node_name] = node.content
  end
  result << hash
end
puts result.inspect

有关详细信息，请参阅文档：http://nokogiri.org/Nokogiri/XML/Node.html

相关内容

最新更新

热门标签：