Skip to content
/ gse Public
forked from go-ego/gse

Go efficient text segmentation; support english, chinese, japanese and other. Go 语言高性能分词

License

Notifications You must be signed in to change notification settings

lazytooo/gse

This branch is 465 commits behind go-ego/gse:master.

Folders and files

NameName
Last commit message
Last commit date

Latest commit

d61867d · Jun 28, 2019
Oct 10, 2018
Feb 19, 2019
Mar 17, 2019
Jun 8, 2019
May 28, 2019
Jun 28, 2019
May 31, 2019
Feb 13, 2019
Mar 17, 2019
Mar 19, 2019
Apr 12, 2018
Feb 27, 2019
Feb 1, 2018
Jun 23, 2017
May 28, 2019
May 31, 2019
Jun 8, 2018
Jan 24, 2019
Feb 24, 2019
Feb 24, 2019
Mar 19, 2019
Mar 19, 2019
Jun 8, 2019
Apr 12, 2018
Jan 26, 2019
Mar 11, 2019
Apr 3, 2018
Feb 17, 2019
Jan 15, 2019

Repository files navigation

gse

Go efficient text segmentation; support english, chinese, japanese and other.

CircleCI Status codecov Build Status Go Report Card GoDoc GitHub release Join the chat at https://gitter.im/go-ego/ego

简体中文

Dictionary with double array trie (Double-Array Trie) to achieve, Sender algorithm is the shortest path based on word frequency plus dynamic programming, and DAG and HMM algorithm word segmentation.

Support common, search engine, full mode, precise mode and HMM mode multiple word segmentation modes, support user dictionary, POS tagging, run JSON RPC service.

Support HMM cut text use Viterbi algorithm.

Text Segmentation speed single thread 9.2MB/s,goroutines concurrent 26.8MB/s. HMM text segmentation single thread 3.2MB/s. (2core 4threads Macbook Pro).

Binding:

gse-bind, binding JavaScript and other, support more language.

Install / update

go get -u github.com/go-ego/gse
go get -u github.com/go-ego/re

re gse

To create a new gse application

$ re gse my-gse

re run

To run the application we just created, you can navigate to the application folder and execute:

$ cd my-gse && re run

Use

package main

import (
	"fmt"

	"github.com/go-ego/gse"
)

var (
	text = "你好世界, Hello world."

	seg gse.Segmenter
)

func cut() {
	hmm := seg.Cut(text, true)
	fmt.Println("cut use hmm: ", hmm)

	hmm = seg.CutSearch(text, true)
	fmt.Println("cut search use hmm: ", hmm)

	hmm = seg.CutAll(text)
	fmt.Println("cut all: ", hmm)
}

func segCut() {
	// Text Segmentation
	tb := []byte(text)
	fmt.Println(seg.String(tb, true))

	segments := seg.Segment(tb)

	// Handle word segmentation results
	// Support for normal mode and search mode two participle,
	// see the comments in the code ToString function.
	// The search mode is mainly used to provide search engines
	// with as many keywords as possible
	fmt.Println(gse.ToString(segments, true))
}

func main() {
	// Loading the default dictionary
	seg.LoadDict()
	// Load the dictionary
	// seg.LoadDict("your gopath"+"/src/github.com/go-ego/gse/data/dict/dictionary.txt")

	cut()

	segCut()
}

Look at an custom dictionary example

package main

import (
	"fmt"

	"github.com/go-ego/gse"
)

func main() {
	var seg gse.Segmenter
	seg.LoadDict("zh,testdata/test_dict.txt,testdata/test_dict1.txt")

	text1 := []byte("你好世界, Hello world")
	fmt.Println(seg.String(text1, true))

	segments := seg.Segment(text1)
	fmt.Println(gse.ToString(segments))
}

Look at an Chinese example

Look at an Japanese example

Authors

License

Gse is primarily distributed under the terms of both the MIT license and the Apache License (Version 2.0), thanks for sego and jieba.

About

Go efficient text segmentation; support english, chinese, japanese and other. Go 语言高性能分词

Resources

License

Stars

Watchers

Forks

Packages

No packages published

Languages

  • Go 99.8%
  • HTML 0.2%