Skip to main content

Finding Duplicate records and Deleting Duplicate records in TERADATA

Requirement:
Finding duplicates and removing duplicate records by retaining original record in TERADATA

Suppose I am working in an office and My boss told me to enter the details of a person who entered in to office. I have below table structure.
Create Table DUP_EXAMPLE
(
PERSON_NAME VARCHAR2(50),
PERSON_AGE INTEGER,
ADDRS VARCHAR2(150),
PURPOSE VARCHAR2(250),
ENTERED_DATE DATE
)

If a person enters more than once then I have to insert his details more than once.
First time, I inserted below records.

INSERT INTO DUP_EXAMPLE VALUES('Krishna reddy','25','BANGALORE','GENERAL',TO_DATE('01-JAN-2014','DD-MON-YYYY'))
INSERT INTO DUP_EXAMPLE VALUES('Anirudh Allika','25','HYDERABAD','GENERAL',TO_DATE('01-JAN-2014','DD-MON-YYYY'))
INSERT INTO DUP_EXAMPLE VALUES('Ashok Vunnam','25','CHENNAI','INTERVIEW',TO_DATE('01-JAN-2014','DD-MON-YYYY'))

And on same day the person named Ashok came again to office and I entered once again into table.
INSERT INTO DUP_EXAMPLE VALUES ('Ashok Vunnam','25','CHENNAI','INTERVIEW',TO_DATE('01-JAN-2014','DD-MON-YYYY'))

Now, I have below data in the table.
SELECT * FROM DUP_EXAMPLE

PERSON_NAME
PERSON_AGE
ADDRS
PURPOSE
ENTERED_DATE
Krishna reddy
25
BANGALORE
GENERAL
01-JAN-2014
Anirudh Allika
25
HYDERABAD
GENERAL
01-JAN-2014
Ashok Vunnam
25
CHENNAI
INTERVIEW
01-JAN-2014
Ashok Vunnam
25
CHENNAI
INTERVIEW
01-JAN-2014


I have a requirement to get the person details that who entered more than once in a day. So, now I have to run below query to get correct result set.
We can write this query in two ways.
1) First Option:

SELECT
PERSON_NAME,
PERSON_AGE,
ADDRS,
PURPOSE,
ENTERED_DATE,
COUNT(*)
FROM DUP_EXAMPLE
GROUP BY 1,2,3,4,5
HAVING COUNT(*)>1

2) Second Option:

SELECT
PERSON_NAME,
PERSON_AGE,
ADDRS,
PURPOSE,
ENTERED_DATE,
ROW_NUMBER() OVER(PARTITION BY PERSON_NAME,PERSON_AGE,ADDRS,PURPOSE,ENTERED_DATE ORDER BY PERSON_NAME,PERSON_AGE,ADDRS,PURPOSE,ENTERED_DATE) AS RECORD_NUMBER
FROM DUP_EXAMPLE
WHERE RECORD_NUMBER > 1

And we can delete duplicate records by retaining original record using below query.

DELETE FROM DUP_EXAMPLE
WHERE ROW_NUMBER() OVER(PARTITION BY PERSON_NAME,PERSON_AGE,ADDRS,PURPOSE,ENTERED_DATE ORDER BY PERSON_NAME,PERSON_AGE,ADDRS,PURPOSE,ENTERED_DATE) > 1

Note: Wherever you go for interview, you will face this question How to find duplicates and how to delete duplicate records by retaining original record.

Comments

  1. Ordered analytical functions are not allowed in WHERE Clause anymore in teradata

    ReplyDelete
  2. This comment has been removed by the author.

    ReplyDelete

Post a Comment

Popular posts from this blog

Comparing Objects in Informatica

We might face a scenario where there may be difference between PRODUCTION v/s SIT version of code or any environment or between different folders in same environment. In here we go for comparison of objects we can compare between mappings,sessions,workflows In Designer it would be present under "Mappings" tab we can find "Compare" option. In workflow manger under "Tasks & Workfows" tab we can find "Compare" option for tasks and workflows comparison respectively. However the easiest and probably the best practice would be by doing using Repository Manager.In Repository Manager under "Edit" tab we can find "Compare" option. The advantage of using Repository manager it compares all the objects at one go i.e. workflow,session and mapping. Hence reducing the effort of individually checking the mapping and session separately. Once we select the folder and corresponding workflow we Can click compare for checking out ...

Types of Joins in Oracle/Teradata

In Data warehousing, irrespective of schema (snow flake schema or star schema) we are using, we should join dimension and fact tables to analyze the business. Below are the frequently used joins: Inner join Left outer Join Right outer Join Cross join Inner Join: Inner join will give you the matching rows from both the tables. If the join condition is not matching then zero records will return. We should use ON keyword to give join condition. Example: Table1: ID Name 1 Krishna 2 Anirudh 4 Ashok Table2: ID Location 1 Bangalore 3 Chennai 4 Chennai We can join above two tables using inner join based on key column ID. SELECT T1.ID, T1.Name, T2.Location FROM Table1 T1 INNER JOIN Table2 T2 ON T1.ID = T2.ID     If we are using inner join, it will give us matching rows from both the table. Here in this example, we have 2 matching rows i.e. ID 1 and 4. Below will be the result set for the above exa...

Looping using Expression Transformation in Informatica

One of the most common used transformation in Informatica is Expression transformation. In Expression transformation we can perform various operations such as data conversions i.e to_date,to_char, string manipulation such as substr,instr etc. Now coming to one of the widely and prominent task which we perform using Expression transformation is looping a value. Expression transformation has three types of ports i.e. input,variable and output.Only output port values can be propagated to next transformations. So in order to pass values of input and variable ports to next level of transformation these must be assigned to output ports.The order of execution in Expression transformation is top to bottom and first input then variable and finally output ports are processed. let us consider the following scenario   The files should be generated with employee name as file name and that particular file should have the details of that respective employee only, if the employee has more t...