pyspark program for nested loop

2019-06-08 15:33发布

站内文章 / Python

26 0

啃猪蹄的小仙女

女 | 书童

私信

可以将文章内容翻译成中文,广告屏蔽插件可能会导致该功能失效(如失效，请关闭广告屏蔽插件后再试):

问题:

I am new to PySpark and I am trying to understand how can we write multiple nested for loop in PySpark, rough high level example below. Any help will be appreciated.

for ( i=0;i<10;i++)
   for ( j=0;j<10;j++)
       for ( k=0;k<10;k++)
          { 
           print "i"."j"."k"
}

回答1:

In non distributed setting for-loops are rewritten using foreachcombinator, but due to Spark nature map and flatMap are a better choice:

from __future__ import print_function
a_loop = lambda x: ((x, y) for y in xrange(10))
print_me = lambda ((x, y), z): print("{0}.{1}.{2}".format(x, y, z)))

(sc.
    parallelize(xrange(10)).
    flatMap(a_loop).
    flatMap(a_loop).
    foreach(print_me)

Of using itertools.product:

from itertools import product
sc.parallelize(product(xrange(10), repeat=3)).foreach(print)

标签： python for-loop apache-spark pyspark rdd

啃猪蹄的小仙女

女 | 书童

私信

收藏的人(0)

Ta的文章更多文章

0条评论

还没有人评论过~

pyspark program for nested loop

问题:

回答1:

收藏的人(0)

举报内容

检举类型

检举原因

检举说明(必填)

打开微信“扫一扫”，打开网页后点击屏幕右上角分享按钮