Use this file for repository-wide rules. Human contributors should start with CONTRIBUTING.md.
- Resolve the target branch (the PR base, not merely the checked-out feature branch), then use the branch context skill.
- Read versions from build.sbt and environment.yml. Do not copy version numbers into this shared guide.
- Use the narrowest relevant repository skill: Scala changes, local setup, and code review.
- For requests to make one or more issues or PRs "5/5", "200% ready", or merge-ready, use the SynapseML PR loop.
- Never edit
target/; generated files are overwritten. - Never commit or print credentials, keys, connection strings, or
.envfiles. - Do not add RDD-based implementations. Use DataFrame/Dataset APIs so code works with Spark Connect and managed Spark modes.
- Do not rebase or force-push shared port branches.
- Do not add a hand-written
__init__.pymerely to re-export generated classes; a stale__all__can hide public APIs. - Keep existing public JVM signatures and serialized parameter shapes unless a breaking change is explicitly approved.
Ask before changing pipeline.yaml, workflows under
.github/workflows/, release tooling, or dependency
pins. These changes affect every branch and require CI evidence.
| Branch | Use it for |
|---|---|
master |
Ordinary features, fixes, and repository-wide changes |
spark<version> port branches |
Differences required specifically by that Spark port |
Run git branch -r to see which port branches currently exist; this guide does
not name them, so that adding one does not require editing a file that must stay
identical on every branch.
Land cross-version changes on master; port branches receive them by merging
master. When resolving a port-branch merge:
- keep the branch side only for version-driven differences;
- take the
masterside for ordinary fixes; - combine both when each changed the same file for a different reason.
Reachability is not proof that a sync preserved content. Compare the merge base,
master, and the port branch for every conflicted file, and verify that
master additions remain present.
AGENTS.md and CONTRIBUTING.md must stay identical across
branches, so neither may name a Spark, Scala, Java, or Python version, or a path
containing one. Put branch-only facts in the
branch context skill. Every other
document — README.md, the website, and module docs — is free to be
branch- and version-specific, because nothing requires those to match across
branches.
| Module | Path | Purpose |
|---|---|---|
| Core | core/ |
SparkML foundations, IO, codegen, AutoML, causal and exploratory tools |
| Cognitive | cognitive/ |
Azure AI and OpenAI service stages |
| LightGBM | lightgbm/ |
Distributed classifier, regressor, and ranker |
| Vowpal Wabbit | vw/ |
VW integration |
| Deep learning | deep-learning/ |
ONNX Runtime inference |
| OpenCV | opencv/ |
Image transformations |
All modules depend on core; deep-learning also depends on opencv.
Each module follows src/main/scala, optional hand-written
src/main/python, and src/test/{scala,python}. Follow nearby code before
introducing a new pattern.
Public SparkML behavior belongs in Scala. Classes mixing in
Wrappable
generate Python wrappers under
target/scala-*/generated/src/python/synapse/ml/.
For a new or changed stage, verify:
- companion object extends
DefaultParamsReadable[Stage]; - class accepts
uid: Stringand has a random-UID no-arg constructor; Wrappableis present when a Python wrapper is required;SynapseMLLoggingis mixed in,logClass(...)is called, andfit/transformis logged;copy(extra)preserves parameters, normally throughdefaultCopy(extra);transformSchemamatches runtime output;- save/load and generated wrapper behavior are tested.
Hand-written Python may extend a generated _ClassName wrapper when JVM
delegation needs a Python convenience method. Do not move core behavior into
Python.
Use the smallest command that proves the change, then expand when the affected surface requires it. Run SBT with the JDK selected by the local setup skill, not the machine default.
sbt <module>/compile
sbt <module>/Test/compile
sbt "<module>/testOnly fully.qualified.Suite"
sbt <module>/scalastyle <module>/Test/scalastyle
sbt codegen
black --check --extend-exclude 'docs/' .Black is pinned in pyproject.toml; use that version.
Scala tests extend
TestBase.
Add positive, negative, schema, persistence, and end-to-end coverage where the
behavior warrants it. A helper-only test is not proof that the public
transformer path works.
Tests that require Azure credentials should skip when credentials are absent. Do not turn a skip into a pass by embedding a secret. Inspect service tests before running them because some create or delete cloud resources.
- Use a conventional title:
feat:,fix:,test:,docs:,ci:, orchore:. - Target
masterunless the change exists only for a port branch. - Resolve active and suppressed review findings; document why any finding is invalid.
- Trigger Azure validation with
/azp runwhere supported. Branch-specific exceptions are documented in the branch context skill. - Treat GitHub Actions as fast checks and pipeline.yaml as the full build.
- Do not equate green checks with correctness: compare before/after failures, inspect skipped tests, and verify the requested behavior directly.
- Rebase feature PRs onto the latest target before final validation. Merge
masterinto shared port branches instead of rebasing them.
Add only durable, repository-wide guidance that changes an agent's decision. Prefer a link to the source of truth over copied commands, versions, or long examples. Put implementation detail beside the code, contributor process in CONTRIBUTING.md, and branch-specific facts in the branch context skill.