Python中的WordCount
本节包括运行用Python 编写的WordCount示例的源代码和说明。
以下是映射器的源代码。 将代码保存到名为 "
WordCount_mapper.py
的文件中:#!/usr/bin/env python
import sys
for line in sys.stdin:
(key,value) = line.split('\t')
wordList = value.split(' ')
for word in wordList:
print(word.strip() + "\t1")
以下是减速器的源代码。 将代码保存到名为 "
WordCount_reducer.py
的文件中:#!/usr/bin/env python
import sys
word = ""
sum = 0
for line in sys.stdin:
(key,value) = line.split('\t')
if word == key:
sum = sum + int(value)
else:
if word != "":
print(word + "\t" + str(sum))
word = key
sum = 1
if key != "": print(key + "\t" + str(sum))
要针对存储在 "
mapreduce_db
数据库中words1表(包括 "id
和 "sentence
列)中的数据运行此示例,请运行以下命令:mapreduce jar /nz/export/ae/products/netezza/mapreduce/current/mapreducestreaming.
jar
-db mapreduce_db
-input 'words1' 'id' 'sentence'
-output 'results' 'word' 'count'
-mapper 'WordCount_mapper.py'
-mapper_out_key_size 20
-mapper_out_value_size 20
-file <path to WordCount_mapper.py file>
-reducer 'WordCount_reducer.py'
-reducer_out_key_size 20
-reducer_out_value size 20
-file <path to WordCount_reducer.py file>
执行命令后,将创建一个包含两列的 "results
表: word
和 "count
,均为 VARCHAR(20) 类型。