xmlparser

package module
v1.0.1 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: May 21, 2019 License: BSD-3-Clause Imports: 2 Imported by: 0

README

XML Stream Parser for GO

xml-stream-parser is a GO library to parse xml files. It is written to addres the performance issue in default xml package.

Install

go get -u github.com/tamerh/xml-stream-parser

Usage

Let say you have following xml and you want to loop over book as a stream and parse various elements and attributes

<?xml version="1.0" encoding="UTF-8"?>
<bookstore>
   <book ISBN="10-000000-001">
      <title>The Iliad and The Odyssey</title>
      <price>12.95</price>
      <comments>
         <userComment rating="4">Best translation I've read.</userComment>
         <userComment rating="2">I like other versions better.</userComment>
      </comments>
      <description>Homer's two epics of the ancient world, The Iliad & The Odyssey, tell stories as riveting today as when they were written between the eighth and ninth century B.C.</description>
   </book>
   <book ISBN="10-000000-999">
      <title>Anthology of World Literature</title>
      <price>24.95</price>
      <comments>
         <userComment rating="3">Needs more modern literature.</userComment>
         <userComment rating="4">Excellent overview of world literature.</userComment>
      </comments>
      <description>The anthology includes epic and lyric poetry, drama, and prose narrative, with many complete works and a focus on the most influential pieces and authors from each region and time period.</description>
   </book>
</bookstore>

you can use the library like so

//First open your file and create reader. You can also use gzip file check tests
file, _ := os.Open("books2.xml")
defer file.Close()
br := bufio.NewReader(file)

// then create  following channel to read your parsed data from.
var resultChannel = make(chan XMLEntry)

// init parser
var parser = XMLParser{
R:          br, 
// define tag to loop over
LoopTag:    "book",
OutChannel: &resultChannel,
// you can skip tags that you are not interested it relatively speeds up the process
SkipTags:   []string{"description"}, 
}

// start parsing with a go routine
go parser.Parse()

// and finally read parsed data 
for book := range resultChannel {
// print ISBN value
isbn := book.Attrs["ISBN"]
fmt.Println(isbn)

// print title
title := book.Elements["title"][0].InnerText
fmt.Println(title)

// print a user commet which has rating 4
// basically you can walk on all the sub nodes if you have
for _, userComments := range book.Elements["comments"][0].Childs {
	for _, comment := range userComments {
		if comment.Attrs["rating"] == "4" {
			  // print the user comment
			  fmt.Println(comment.InnerText)
			}
		}
	}
}

Documentation

Index

Constants

This section is empty.

Variables

This section is empty.

Functions

This section is empty.

Types

type XMLElement

type XMLElement struct {
	Attrs     map[string]string
	InnerText string
	Childs    map[string][]XMLElement
}

XMLElement is a typical xml elements which keeps the parsed data

type XMLEntry

type XMLEntry struct {
	Attrs    map[string]string
	Elements map[string][]XMLElement
}

XMLEntry is a result of each parsed loop

type XMLParser

type XMLParser struct {
	R             *bufio.Reader
	LoopTag       string
	OutChannel    *chan XMLEntry
	SkipTags      []string
	FinishMessage string
	// contains filtered or unexported fields
}

XMLParser parser/scrapper of xml file For more improvment following can be done 1- skip tags inside the element for now a tag can be skipped only root element of looptag 2- to make it more parallel maybe first just get the looptag content and send it for processing. 3- change slices size and append if applicabale.

func (*XMLParser) Parse

func (x *XMLParser) Parse()

Parse starts parsing xml document

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL