99热这里只有精品5,久久精品国产精品亚洲综合,a天堂中文在线

本文實例講述了Python找出文件中使用率最高的漢字的方法。分享給大家供大家參考。具體分析如下：

這是我初學Python時寫的，為了簡便，我并沒在排序完后再去掉非中文字符，稍微會影響性能（大約增加了25％的時間）。

									# -*- coding: gbk -*- 

									import codecs 

									from time import time 

									from operator import itemgetter 

									def top_words(filename, size=10, encoding='gbk'): 

									  count = {} 

									  for line in codecs.open(filename, 'r', encoding): 

									    for word in line: 

									      if u'\u4E00' <= word <= u'\u9FA5' or u'\uF900' <= word <= u'\uFA2D': 

									        count[word] = 1 + count.get(word, 0) 

									  top_words = sorted(count.iteritems(), key=itemgetter(1), reverse=True)[:size] 

									  print '\n'.join([u'%s : %s次' % (word, times) for word, times in top_words]) 

									begin = time() 

									top_words('空之境界.txt') 

									print '一共耗時 : %s秒' % (time()-begin)

如果想用上新方法，以及讓join的可讀性更高的話，這樣也是可以的：

									# -*- coding: gbk -*- 

									import codecs 

									from time import time 

									from operator import itemgetter 

									from heapq import nlargest 

									def top_words(filename, size=10, encoding='gbk'): 

									  count = {} 

									  for line in codecs.open(filename, 'r', encoding): 

									    for word in line: 

									      if u'\u4E00' <= word <= u'\u9FA5' or u'\uF900' <= word <= u'\uFA2D': 

									        count[word] = 1 + count.get(word, 0) 

									  top_words = nlargest(size, count.iteritems(), key=itemgetter(1)) 

									  for word, times in top_words: 

									    print u'%s : %s次' % (word, times) 

									begin = time() 

									top_words('空之境界.txt') 

									print '一共耗時 : %s秒' % (time()-begin)

或者讓行數更少（好囧的列表綜合）：

									# -*- coding: gbk -*- 

									import codecs 

									from time import time 

									from operator import itemgetter 

									def top_words(filename, size=10, encoding='gbk'): 

									  count = {} 

									  for word in [word for word in codecs.open(filename, 'r', encoding).read() if u'\u4E00' <= word <= u'\u9FA5' or u'\uF900' <= word <= u'\uFA2D']: 

									    count[word] = 1 + count.get(word, 0) 

									  top_words = sorted(count.iteritems(), key=itemgetter(1), reverse=True)[:size] 

									  print '\n'.join([u'%s : %s次' % (word, times) for word, times in top_words]) 

									begin = time() 

									top_words('空之境界.txt') 

									print '一共耗時 : %s秒' % (time()-begin)

此外還可以引入with語句，這樣只需一行就能獲得異常安全性。
3者性能幾乎一樣，結果如下：

									的 : 17533次

									是 : 8581次

									不 : 6375次

									我 : 6168次

									了 : 5586次

									一 : 5197次

									這 : 4394次

									在 : 4264次

									有 : 4188次

									人 : 4025次

									一共耗時 : 0.5秒

引入psyco模塊的成績：

									的 : 17533次

									是 : 8581次

									不 : 6375次

									我 : 6168次

									了 : 5586次

									一 : 5197次

									這 : 4394次

									在 : 4264次

									有 : 4188次

									人 : 4025次

									一共耗時 : 0.280999898911秒