Introduction
Let's say the points (x1, y1) and (x2, y2) are points in 2-dimensional space and distance by using the Pythagorean formula like below.

Then, the Euclidean distance between (x1, y1) and (x2, y2) is,
d = √(x2 - x1)^2 + (y2 - y1) ^2
So, in short form,

- Training Data
This set of data contains all the information included with classifications like below.
This training data includes classification with given x, y values.

- Test Data
This set of data contains only the values of x and y. Its classification type would be predicted based on the training data.
This set of training data doesn't contain a classification type. So, it will be predicted.
Implementation

Import the below libraries.
- import csv
- import sys
- from collections import Counter
- from math import sqrt
- x = []
- y = []
- z = []
- with open('training_data.csv','rt') as f:
- reader = csv.reader(f)
- for row in reader:
- x.append(float(row[0]))
- y.append(float(row[1]))
- z.append(row[2])
- coordinates = list(zip(x,y))
- input_data = {coordinates[i]:z[i] for i in range(len(coordinates))}
- test_x = []
- test_y = []
- with open('test_data.csv', 'rt') as f:
- reader = csv.reader(f)
- for row in reader:
- test_x.append(float(row[0]))
- test_y.append(float(row[1]))
- test_coordinates = list(zip(test_x, test_y))
- print (test_coordinates)
- def euclidean_distance(x, y):
- if len(x) != len(y):
- return "Error: try equal length vectors"
- else:
- return sqrt(sum([(x[i]-y[i])**2 for i in range(len(y))]))
- def knn_classifier(neighbors, input_data):
- knn = [input_data[i] for i in neighbors]
- knn = Counter(knn)
- classifier, _ = knn.most_common(1)[0]
- return classifier
- def neighbors(k, trained_points, new_point):
- neighbor_distances = {}
- for point in trained_points:
- if point not in neighbor_distances:
- neighbor_distances[point] = euclidean_distance(point, new_point)
- least_common = sorted(neighbor_distances.items(), key = lambda x: x[1])
- k_nearest_neighbors = list(zip(*least_common[:k]))
- return list(k_nearest_neighbors[0])
- results = {}
- for item in test_coordinates:
- results[item] = knn_classifier(neighbors(3,input_data.keys(), item), input_data)
- print (results)

Here, x and y data have been classified into different groups. I have attached the zipped Python code. Python 3 or above will be required to execute this code.
Join the conversation! Your thoughts help the community grow.