有效地计算Pyspark中的连接组件

我正在尝试找到城市中朋友的连接组件。我的数据是带有城市属性的边缘列表。

城市|src |dest

休斯顿凯尔 ->本尼

休斯顿本尼 ->查尔斯

休斯顿查尔斯 ->丹尼

奥马哈颂歌 -> Brian

等。

我知道Pyspark的GraphX库的连接组件功能将在图的所有边缘上迭代以找到连接的组件，我想避免使用。我该怎么办？

编辑：我以为我可以做

之类的事情

从dataframe中选择Connected_components（*）Groupby City

connected_components生成项目列表。

假设您的数据就像这样

import org.apache.spark._
import org.graphframes._
val l = List(("Houston","Kyle","Benny"),("Houston","Benny","charles"),
            ("Houston","Charles","Denny"),("Omaha","carol","Brian"),
            ("Omaha","Brian","Daniel"),("Omaha","Sara","Marry"))
var df = spark.createDataFrame(l).toDF("city","src","dst")

创建要运行连接组件的城市列表 cities = List("Houston","Omaha")

现在，在城市列表中的每个城市的城市列上运行一个过滤器，然后从结果框架中创建一个边缘和顶点数据框。从这些边缘和顶点dataframes创建一个GraphFrame，并运行连接的组件算法

val cities = List("Houston","Omaha")
for(city <- cities){
    val edges = df.filter(df("city") === city).drop("city")
    val vert = edges.select("src").union(edges.select("dst")).
                     distinct.select(col("src").alias("id"))
    val g = GraphFrame(vert,edges)
    val res = g.connectedComponents.run()
    res.select("id", "component").orderBy("component").show()
}

输出

|     id|   component|
+-------+------------+
|   Kyle|249108103168|
|charles|249108103168|
|  Benny|249108103168|
|Charles|721554505728|
|  Denny|721554505728|
+-------+------------+
+------+------------+                                                           
|    id|   component|
+------+------------+
| Marry|858993459200|
|  Sara|858993459200|
| Brian|944892805120|
| carol|944892805120|
|Daniel|944892805120|
+------+------------+

相关内容

最新更新

热门标签：