2026-07-15-workhours-data-cleanup.md 15 KB

Workhours Data Cleanup Implementation Plan

For agentic workers: REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (- [ ]) syntax for tracking.

Goal: Safely identify and delete duplicate required-hours records plus records after an employee's offline date, while preserving all other empty or unmatched data.

Architecture: Add a package-private pure resolver that deterministically selects one keeper per employee/date key. WorkHoursCalcService will scan required-hours records by month, exclude post-offline records from duplicate grouping so the two deletion sets are disjoint, and expose a dry-run-first controller endpoint. Production deletion remains blocked until online dry-run counts are reported and the user confirms a second time.

Tech Stack: Java 8, Spring Boot 2.1, JUnit 4, Mockito, DingTalk YiDa form APIs, Maven.

2026-07-15 实际执行修订

用户确认实际顺序调整为:隔离分支本地启动服务处理正式宜搭数据 → 更新文档归档 → 最后整理提交并部署。没有为了 dry-run 先部署生产服务。

当前结果:

  • 重复记录删除 35 条,失败 0。
  • 离职日后记录删除 796 条,失败 0。
  • 删除后重复、离职后、未来日期三项 dry-run 的 toDelete 均为 0。
  • 两份业务文档已更新;最终生产部署仍待执行。

Task 1: Deterministic Duplicate Resolver

Files:

  • Create: mjava-akdsbeisen/src/main/java/com/malk/service/workhours/WorkHoursDuplicateResolver.java
  • Create: mjava-akdsbeisen/src/test/java/com/malk/service/workhours/WorkHoursDuplicateResolverTest.java

  • [ ] Step 1: Write the failing resolver tests

Create tests for completeness, modified-time tie-breaking, and invalid keys:

@Test
public void resolveShouldKeepMostCompleteCandidate() {
    Candidate sparse = candidate("old", "u1|2026-07-15", 2, 200L, 200L);
    Candidate complete = candidate("complete", "u1|2026-07-15", 5, 100L, 100L);

    Resolution result = WorkHoursDuplicateResolver.resolve(Arrays.asList(sparse, complete));

    assertEquals(Collections.singletonList("old"), result.getDeleteInstanceIds());
    assertEquals("complete", result.getGroups().get(0).getKeepInstanceId());
}

@Test
public void resolveShouldKeepMostRecentlyModifiedWhenCompletenessMatches() {
    Candidate older = candidate("older", "u1|2026-07-15", 5, 100L, 100L);
    Candidate newer = candidate("newer", "u1|2026-07-15", 5, 200L, 100L);

    Resolution result = WorkHoursDuplicateResolver.resolve(Arrays.asList(older, newer));

    assertEquals(Collections.singletonList("older"), result.getDeleteInstanceIds());
}

@Test
public void resolveShouldPreserveCandidatesWithoutReliableKey() {
    Resolution result = WorkHoursDuplicateResolver.resolve(Collections.singletonList(
            candidate("invalid", null, 5, 200L, 100L)));

    assertEquals(1, result.getSkippedInvalidKey());
    assertTrue(result.getDeleteInstanceIds().isEmpty());
}
  • Step 2: Run the resolver tests and verify RED

Run:

mvn -pl mjava-akdsbeisen -am \
  -Dmaven.test.skip=false -DskipTests=false \
  -Dsurefire.failIfNoSpecifiedTests=false \
  -Dtest=WorkHoursDuplicateResolverTest test

Expected: test compilation fails because WorkHoursDuplicateResolver does not exist.

  • Step 3: Implement the minimal resolver

Create package-private typed classes:

final class WorkHoursDuplicateResolver {
    private WorkHoursDuplicateResolver() {
    }

    static Resolution resolve(List<Candidate> candidates) {
        Map<String, List<Candidate>> grouped = new LinkedHashMap<>();
        int skippedInvalidKey = 0;
        for (Candidate candidate : candidates) {
            if (candidate.getKey() == null || candidate.getKey().isEmpty()) {
                skippedInvalidKey++;
                continue;
            }
            grouped.computeIfAbsent(candidate.getKey(), key -> new ArrayList<>()).add(candidate);
        }

        List<DuplicateGroup> groups = new ArrayList<>();
        List<String> deleteIds = new ArrayList<>();
        for (Map.Entry<String, List<Candidate>> entry : grouped.entrySet()) {
            List<Candidate> group = entry.getValue();
            if (group.size() <= 1) continue;
            group.sort(KEEPER_ORDER);
            Candidate keeper = group.get(0);
            List<String> groupDeleteIds = group.subList(1, group.size()).stream()
                    .map(Candidate::getInstanceId)
                    .collect(Collectors.toList());
            deleteIds.addAll(groupDeleteIds);
            groups.add(new DuplicateGroup(entry.getKey(), keeper.getInstanceId(), groupDeleteIds));
        }
        return new Resolution(grouped.size(), skippedInvalidKey, groups, deleteIds);
    }
}

KEEPER_ORDER sorts completeness, modified time, and created time descending, then instance ID ascending. Add explicit constructors and typed getters for Candidate, DuplicateGroup, and Resolution because the project targets Java 8.

  • Step 4: Run the resolver tests and verify GREEN

Run the command from Step 2.

Expected: 3 tests, 0 failures, 0 errors.

  • Step 5: Commit the resolver
git add mjava-akdsbeisen/src/main/java/com/malk/service/workhours/WorkHoursDuplicateResolver.java \
        mjava-akdsbeisen/src/test/java/com/malk/service/workhours/WorkHoursDuplicateResolverTest.java
git commit -m "feat(workhours): 增加重复记录保留规则"

Task 2: Lock the Offline-Date Boundary

Files:

  • Modify: mjava-akdsbeisen/src/main/java/com/malk/service/workhours/WorkHoursCalcService.java
  • Modify: mjava-akdsbeisen/src/test/java/com/malk/service/workhours/WorkHoursCalcServiceTest.java

  • [ ] Step 1: Write a failing boundary test

@Test
public void isAfterOfflineDateShouldKeepOfflineDayAndRejectFollowingDay() {
    LocalDate offlineDate = LocalDate.of(2026, 7, 15);

    assertFalse(WorkHoursCalcService.isAfterOfflineDate(offlineDate, offlineDate));
    assertTrue(WorkHoursCalcService.isAfterOfflineDate(offlineDate.plusDays(1), offlineDate));
    assertFalse(WorkHoursCalcService.isAfterOfflineDate(offlineDate.plusDays(1), null));
}
  • Step 2: Run the service test and verify RED

Run:

mvn -pl mjava-akdsbeisen -am \
  -Dmaven.test.skip=false -DskipTests=false \
  -Dsurefire.failIfNoSpecifiedTests=false \
  -Dtest=WorkHoursCalcServiceTest test

Expected: test compilation fails because isAfterOfflineDate does not exist.

  • Step 3: Implement and reuse the boundary helper
static boolean isAfterOfflineDate(LocalDate workDay, LocalDate offlineDate) {
    return workDay != null && offlineDate != null && workDay.isAfter(offlineDate);
}

Replace the three direct workDay.isAfter(offlineDate) checks in single-day sync, concurrent upsert, and offline cleanup with this helper.

  • Step 4: Run the service test and verify GREEN

Run the command from Step 2.

Expected: all WorkHoursCalcServiceTest tests pass.

  • Step 5: Commit the boundary test and refactor
git add mjava-akdsbeisen/src/main/java/com/malk/service/workhours/WorkHoursCalcService.java \
        mjava-akdsbeisen/src/test/java/com/malk/service/workhours/WorkHoursCalcServiceTest.java
git commit -m "test(workhours): 锁定离职日期过滤边界"

Task 3: Dry-Run-First Duplicate Cleanup Endpoint

Files:

  • Modify: mjava-akdsbeisen/src/main/java/com/malk/service/workhours/WorkHoursCalcService.java
  • Modify: mjava-akdsbeisen/src/main/java/com/malk/controller/WorkHoursController.java
  • Modify: mjava-akdsbeisen/src/test/java/com/malk/service/workhours/WorkHoursCalcServiceTest.java

  • [ ] Step 1: Write a failing dry-run integration test

Mock two required-hours records with the same employee/date. Make one candidate more complete, and verify dry-run reports one deletion without invoking YiDa delete:

@Test
public void cleanupDuplicateHoursDryRunShouldReportOneDeletionWithoutDeleting() {
    YDClient ydClient = mock(YDClient.class);
    WHConf conf = requiredHoursConf();
    DDR_New<Object> page = pageOf(
            requiredHoursRecord("sparse", "employee-1", LocalDate.of(2026, 7, 15)),
            completeRequiredHoursRecord("complete", "employee-1", LocalDate.of(2026, 7, 15)));
    when(ydClient.queryData(any(YDParam.class), eq(YDConf.FORM_QUERY.retrieve_search_form)))
            .thenReturn(personnelPageWithoutOfflineDate(), page, emptyPage());
    WorkHoursCalcService service = serviceWith(ydClient, conf);

    Map<String, Object> stats = service.cleanupDuplicateHours(true);

    assertEquals(1, stats.get("duplicateGroups"));
    assertEquals(1, stats.get("toDelete"));
    assertEquals(0, stats.get("deleted"));
    verify(ydClient, never()).operateData(any(YDParam.class), eq(YDConf.FORM_OPERATION.delete_batch));
}
  • Step 2: Run the service test and verify RED

Run the Task 2 test command.

Expected: compilation fails because cleanupDuplicateHours does not exist.

  • Step 3: Implement monthly scan and candidate mapping

Add:

public Map<String, Object> cleanupDuplicateHours(boolean dryRun)

Implementation requirements:

  1. Load all personnel details and build userId -> offlineDate.
  2. Scan required-hours records from 2026-04 through the current month using the existing monthly date-range pattern.
  3. Build candidate key only when employee and work day are both present.
  4. Exclude isAfterOfflineDate(workDay, offlineDate) records from duplicate grouping and count them in excludedAfterOffline; these are handled by cleanupAfterOffline.
  5. Score completeness using hours, Manager, employee number, attribute, department, company, and CF fields.
  6. Normalize gmtModified and gmtCreate to epoch milliseconds when numeric; use 0 for unavailable values.
  7. Resolve duplicates with WorkHoursDuplicateResolver.
  8. In dry-run mode, return stats and up to five group samples without calling delete.
  9. In formal mode, delete resolver IDs in batches of at most 100 and report deleted and fail.
  • Step 4: Add the controller endpoint
@GetMapping("/cleanup-duplicates")
public Map<String, Object> cleanupDuplicates(
        @RequestParam(defaultValue = "true") boolean dryRun) {
    Map<String, Object> result = new LinkedHashMap<>();
    try {
        long start = System.currentTimeMillis();
        Map<String, Object> stats = workHoursCalcService.cleanupDuplicateHours(dryRun);
        result.put("success", true);
        result.put("message", dryRun ? "重复工时清理预览完成(未删除)" : "重复工时清理完成");
        result.put("stats", stats);
        result.put("costMs", System.currentTimeMillis() - start);
    } catch (Exception e) {
        log.error("重复工时清理失败", e);
        result.put("success", false);
        result.put("message", e.getMessage());
    }
    return result;
}

Default dryRun=true is intentional so an omitted parameter cannot delete production data.

  • Step 5: Run focused tests and verify GREEN

Run:

mvn -pl mjava-akdsbeisen -am \
  -Dmaven.test.skip=false -DskipTests=false \
  -Dsurefire.failIfNoSpecifiedTests=false \
  -Dtest=WorkHoursDuplicateResolverTest,WorkHoursCalcServiceTest,WorkHoursTimerScheduleTest test

Expected: all focused tests pass with 0 failures and 0 errors.

  • Step 6: Commit the endpoint
git add mjava-akdsbeisen/src/main/java/com/malk/controller/WorkHoursController.java \
        mjava-akdsbeisen/src/main/java/com/malk/service/workhours/WorkHoursCalcService.java \
        mjava-akdsbeisen/src/test/java/com/malk/service/workhours/WorkHoursCalcServiceTest.java
git commit -m "feat(workhours): 增加重复数据安全清理接口"

Task 4: Verify, Deploy the Dry-Run Capability, and Audit Production

Files:

  • Verify only; no additional source files.

  • [ ] Step 1: Run the final build

mvn -q -pl mjava-akdsbeisen -am clean package \
  -Dmaven.test.skip=false -DskipTests=false \
  -Dsurefire.failIfNoSpecifiedTests=false \
  -Dtest=WorkHoursDuplicateResolverTest,WorkHoursCalcServiceTest,WorkHoursTimerScheduleTest

Expected: exit code 0; all test report XML files show zero failures and errors.

  • Step 2: Verify source and artifact
git diff --check HEAD~3 HEAD
jar tf mjava-akdsbeisen/target/mjava-akdsbeisen.jar | rg '(^|/)h2-[^/]*\.jar$' || true
shasum -a 256 mjava-akdsbeisen/target/mjava-akdsbeisen.jar

Expected: no diff errors, no H2 dependency, and a SHA-256 value.

  • Step 3: Deploy using the already approved protected workflow

Use the isolated JAR, preserve the original dirty repository's local JAR, and run:

/Users/malk/.agents/skills/mcli/deploy/bin/deploy.sh akds --skip-build --backup --yes

Expected: remote backup created, upload completes, and ./server.sh status reports RUNNING with prod profile.

  • Step 4: Run online read-only audits
curl -sS 'http://127.0.0.1:9055/api/akds/workhours/cleanup-duplicates?dryRun=true'
curl -sS 'http://127.0.0.1:9055/api/akds/workhours/cleanup-after-offline?dryRun=true'

Expected: duplicate audit returns exact disjoint duplicate counts; offline audit continues to return 796 before deletion.

  • Step 5: Stop before deletion and request confirmation

Report target form, duplicate groups, duplicate delete count, offline delete count, combined delete count, keeper rule, and samples. Do not call either formal endpoint until the user explicitly confirms.

Task 5: Execute Confirmed Deletion and Verify Idempotency

Files:

  • Modify after successful deletion: /Users/malk/Desktop/Tech/claude/后端/阿科德斯/应填报工时月度计算.md
  • Modify after successful deletion: /Users/malk/Desktop/Tech/claude/临时/阿科德斯-应报工时修复-2026-07-15.md

  • [x] Step 1: Delete valid-record duplicates after confirmation

curl -sS 'http://127.0.0.1:9055/api/akds/workhours/cleanup-duplicates?dryRun=false'

Expected: deleted == toDelete and fail == 0.

  • Step 2: Delete post-offline records
curl -sS 'http://127.0.0.1:9055/api/akds/workhours/cleanup-after-offline?dryRun=false'

Expected: deleted == toDelete and fail == 0.

  • Step 3: Re-run all read-only audits
curl -sS 'http://127.0.0.1:9055/api/akds/workhours/cleanup-duplicates?dryRun=true'
curl -sS 'http://127.0.0.1:9055/api/akds/workhours/cleanup-after-offline?dryRun=true'
curl -sS 'http://127.0.0.1:9055/api/akds/workhours/cleanup-future?dryRun=true'

Expected: every response has toDelete=0 and fail=0.

  • Step 4: Update both workhours documents

Record actual duplicate groups, duplicate deletions, offline deletions, verification results, commit IDs, and deployment state. Keep unmatched personnel and empty-source preservation rules explicit.

  • Step 5: Final repository and remote verification
git status --short
ssh root@120.55.113.155 'cd /home/server/akds && ./server.sh status'

Expected: code worktree clean and remote service RUNNING.